stable-baselines3

mirror of https://github.com/saymrwulf/stable-baselines3.git synced 2026-05-28 22:56:53 +00:00

Author	SHA1	Message	Date
Antonin RAFFIN	494ebfd20a	Hotfix PPO + gSDE (#53 ) * Fix variable being passed with gradients * Update changelog * Bump version * Fixes #54	2020-06-10 18:58:35 +02:00
Anssi	b833207142	Add some missing tests, update VecNormalize and RolloutBuffer (#50 ) * Change saving/loading normalization parameters to use single pickle file * Remove 'use_gae' from RolloutBuffer compute_returns function * Add some missing tests for normalizer, nan-checker and PPO clip_value_fn argument * Update changelog * Fix typo * Use proper pytest.raises for catching errors in tests * Add comment on GAE and how to obtain non-GAE behaviour * Remove save/load_running_average from VecNormalize in favor of load/save * Update changelog * Update docstring * Add accidentally removed tests for VecNormalize Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-06-10 12:09:04 +02:00
Anssi	44f8218df0	Review of code (A2C, PPO and refactoring) (#35 ) * Split torch module code into torch_layers file * Updated reference to CNN * Change 'CxWxH' to 'CxHxW', as per common notion * Fix missing import in policies.py * Move PPOPolicy to OnlineActorCriticPolicy * Create OnPolicyRLModel from PPO, and make A2C and PPO inherit * Update A2C optimizer comment * Clean weight init scales for clarity * Fix A2C log_interval default parameter * Rename 'progress' to 'progress_remaining * Rename 'Models' to 'Algorithms' * Rename 'OnlineActorCriticPolicy' to 'ActorCriticPolicy' * Move static functions out from BaseAlgorithm * Move on/off_policy base algorithms to their own files * Add files for A2C/PPO * Fix docs * Fix pytype * Update documentation on OnPolicyAlgorithm * Add proper doctstring for on_policy rollout gathering * Add bit clarification on the mlppolicy/cnnpolicy naming * Move static function is_vectorized_policies to utils.py * Checking docstrings, pep8 fixes * Update changelog * Clean changelog * Remove policy warnings for sac/td3 * Add monitor_wrapper for OnPolicyAlgorithm. Clean tb logging variables. Add parameter keywords to OffPolicyAlgorithm super init Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-06-09 13:54:18 +02:00
Antonin RAFFIN	11d33eb4ae	Fix gSDE loading issue in test mode (#45 ) * Fix gSDE loading issue in test mode * Forward `reset_noise` method * Re-add `make_actor` * Reformat	2020-06-08 11:15:10 +02:00
Antonin RAFFIN	353ea81080	Fix several VecEnv issues, add `fork` start method to tests (#43 ) * Fix several VecEnv issues, add `fork` start method to tests * Fix signature	2020-06-04 11:22:12 +02:00
Antonin RAFFIN	403fff5d50	Pre-Release v0.6.0 (#39 ) * Prepare release * Update docker images	2020-06-01 13:09:47 +02:00
Roland Gavrilescu	bb01253261	Tensorboard integration (#30 ) * init commit tensorboard-integration * Added tb logger to ppo (with output exclusions) * fixed truncated stdout * categorize stdout outputs by tag * separated exclusions from values, added missing logs * saving exclusions as dict instead of list * reformatting, auto run indexing * included renaming suggestions, fixed tests * tb support for sac * linting * moved logging to base class * tb support for td3 * removed histograms, non-verbose output working * modifed changelog * linting * fixed type error * moved logger config to utils * removed episode_rewards log from ppo * Enable tensorboard in tests * Remove unused import * Update logger sub titles * Minor edit for PPO * Update logger and tb log folder * Pass correct logger to Callbacks * updated docs * added tb example image to docs * add support for continuing training in tensorboard * added tensorboard to docs index * added tb test * moved logger config to _setup_learn, updated tests * accessing verbose from base class * Update doc and tests * Rename session -> time * Update version * Update logger truncate * Update types * Remove duplicated code Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-06-01 11:55:44 +02:00
mloo3	42f432c79c	Fix TD3 Example Code Documentation (#38 ) Fix TD3's example code	2020-06-01 11:37:42 +03:00
Stelios Tymvios	78e8d405d7	Implemented Vectorized Action Noise (#34 ) * Implemented Vectorized Action Noise Vectorized Action Noise allows for multiple instances of ActionNoiseProcesses to run in parallel. This makes it easier to run TD3/SAC/DDPG with VecEnv. * fixed linting issues * make test function name consistent Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * sanity checks and more detailed test * Update stable_baselines3/common/noise.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Added assertion error message in noises setter * Corrected tests to reflect change to AssertionError from ValueError Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-05-27 09:53:01 +02:00
Antonin RAFFIN	9b42b9717a	Fix `sde_sample_freq` for SAC (#32 ) * Fix `sde_sample_freq` for SAC * [ci skip] Add Acknowledgments	2020-05-24 16:44:44 +02:00
Tarik Kelestemur	b1322ff5d6	Fix cmd_util.py imports (#24 ) * fix cmd_util.py imports * Update changelog.rst Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-05-19 10:19:16 +02:00
Roland Gavrilescu	91adefdb4b	Support for MultiBinary / MultiDiscrete spaces (#13 ) * multicategorical dist and test * fixed List annotation * bernoulli dist and test * added distributions to preprocessing (needs testing) * fixed and tested distributions * added changelog and fixed ppo policy * minor fix * dist fixes, added test_spaces * clean up * modified changelog * additional fixes * minor changelog mod * hot encoding fix, flake8 clean up * lint tests * preprocessing fix * fixed bernoulli bug * removed commented prints * Update changelog.rst * included suggested modifications * linting fix * increased space dim * Update doc and tests Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-05-18 14:42:13 +02:00
Antonin RAFFIN	15ff6d47ee	Documentation update and style fixes (#21 ) * Update doc: add gSDE * Fix codestyle * Remove travis script * Add lint check to gitlab	2020-05-15 13:54:06 +02:00
Antonin RAFFIN	54f6f5b6fb	Add flake8 linter and Github CI (#19 ) * Cleanup code * Add flake8 lint and github workflow * Update build matrix * Relax precision for python3.7	2020-05-12 17:55:01 +02:00
Antonin RAFFIN	b02afd6ee3	Doc update (#15 )	2020-05-11 12:28:43 +02:00
Antonin RAFFIN	257a40ef4b	Add Gitlab CI (#12 ) * Test gitlab-ci * Try different image * Add pytest and doc build * Fix command * Fix image used for CI * Seperate pytest builds * Fix weird seg fault in docker image due to FakeImageEnv * Fix make command * [ci skip] Add space in the badges * Fix CI failures * Re-install opencv * Use opencv-headless * Test with new docker image	2020-05-09 23:10:49 +02:00
Kinal Mehta	b1f5db1bb2	Add CONTRIBUTION.md link in README (#2 ) * Fix CONTRIBUTION.md link in README * Update changelog.rst Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org>	2020-05-09 13:01:15 +02:00
Antonin RAFFIN	c20af230f7	Remove SDE support for TD3	2020-05-08 15:00:34 +02:00
Antonin RAFFIN	97aea21349	Update minimum gym version	2020-05-08 12:43:42 +02:00
Antonin RAFFIN	e6ff4bbd6c	Update setup	2020-05-07 16:24:19 +02:00
Antonin RAFFIN	aa66012764	Update requirements	2020-05-07 16:21:33 +02:00
Antonin RAFFIN	8046a24719	More doc + sync VecEnvs + atari	2020-05-07 16:08:23 +02:00
Antonin RAFFIN	73afaf157c	Add version.txt to package	2020-05-07 12:19:29 +02:00
Antonin RAFFIN	d17f29c8ad	Add base doc	2020-05-07 10:10:51 +02:00
Antonin RAFFIN	580317158b	Update changelog	2020-05-05 17:21:56 +02:00
Antonin RAFFIN	0481fbe727	Update changelog	2020-05-05 16:54:33 +02:00
Antonin RAFFIN	2c34a4d694	Sync with Stable-Baselines	2020-05-05 16:28:38 +02:00
Antonin RAFFIN	d542732c8d	Rename to stable-baselines3	2020-05-05 15:02:35 +02:00
Antonin RAFFIN	88cee2ba55	Add type hints and f-strings to logger	2020-05-05 14:49:32 +02:00
Antonin RAFFIN	041f2bc59a	Cleanup, bug fixes + more tests	2020-04-22 13:14:22 +02:00
Antonin RAFFIN	8aac9e819d	Add `VecTransposeImage` and fix for SAC	2020-04-21 20:41:58 +02:00
Antonin RAFFIN	93c2a01f91	Start CNN support (failing for SAC)	2020-04-21 16:22:46 +02:00
Antonin RAFFIN	f347474e6a	Independent save/load for policies	2020-04-20 15:59:44 +02:00
Antonin RAFFIN	17f9246257	Add `get_device` util and fix `squash_output`	2020-04-20 15:43:11 +02:00
Antonin RAFFIN	aa1026ee87	Added ``optimizer` `and` `optimizer_kwargs` `to` `policy_kwargs``	2020-04-17 15:13:45 +02:00
Antonin RAFFIN	0e44cdce44	Fixed ``reset_num_timesteps`` behavior	2020-04-17 12:36:27 +02:00
Antonin RAFFIN	08a22c4834	Release 0.4.0	2020-04-14 18:13:51 +02:00
Antonin RAFFIN	fdecd512db	Add save/load weights for policies and refactor action distributions	2020-03-31 16:29:13 +02:00
Antonin RAFFIN	fa599c65a6	Add support for Discrete observation spaces	2020-03-25 16:42:05 +01:00
Antonin RAFFIN	72a88a8d92	Fix type hint for activation fn	2020-03-24 10:10:37 +01:00
Antonin RAFFIN	ba18258af6	Add proper preprocessing	2020-03-23 17:15:30 +01:00
Antonin RAFFIN	dcb54b5301	Remove CEMRL	2020-03-23 14:48:38 +01:00
Antonin RAFFIN	57b37513b6	Refactor handling of obs and action space + remove duplicated code	2020-03-20 10:09:09 +01:00
Antonin RAFFIN	7251b9d2c2	Release v0.3.0	2020-03-19 11:11:36 +01:00
Antonin RAFFIN	fd9e73cfb8	Fix entropy computation	2020-03-19 10:19:48 +01:00
Antonin RAFFIN	9485b90a41	Sync predict with SB and add version file	2020-03-18 15:11:19 +01:00
Antonin RAFFIN	c3187604bc	Code cleanup: rename lr to lr_schedule + typing	2020-03-16 14:01:32 +01:00
Antonin Raffin	29d7018265	Add better logging for SAC and PPO	2020-03-13 11:43:12 +01:00
Antonin Raffin	c39421fa64	Fix colors in results plotter	2020-03-13 10:59:16 +01:00
Antonin Raffin	b64873ffff	Sync callbacks	2020-03-12 12:34:25 +01:00

1 2

73 commits