stable-baselines3/docs/guide/vec_envs.rst

.. _vec_env:

.. automodule:: stable_baselines3.common.vec_env

Vectorized Environments
=======================

Vectorized Environments are a method for stacking multiple independent environments into a single environment.
Instead of training an RL agent on 1 environment per step, it allows us to train it on ``n`` environments per step.
Because of this, ``actions`` passed to the environment are now a vector (of dimension ``n``).
It is the same for ``observations``, ``rewards`` and end of episode signals (``dones``).
In the case of non-array observation spaces such as ``Dict`` or ``Tuple``, where different sub-spaces
may have different shapes, the sub-observations are vectors (of dimension ``n``).

============= ======= ============ ======== ========= ================
Name          ``Box`` ``Discrete`` ``Dict`` ``Tuple`` Multi Processing
============= ======= ============ ======== ========= ================
DummyVecEnv   ✔️       ✔️           ✔️        ✔️         ❌️
SubprocVecEnv ✔️       ✔️           ✔️        ✔️         ✔️
============= ======= ============ ======== ========= ================

.. note::

	Vectorized environments are required when using wrappers for frame-stacking or normalization.

.. note::

	When using vectorized environments, the environments are automatically reset at the end of each episode.
	Thus, the observation returned for the i-th environment when ``done[i]`` is true will in fact be the first observation of the next episode, not the last observation of the episode that has just terminated.
	You can access the "real" final observation of the terminated episode—that is, the one that accompanied the ``done`` event provided by the underlying environment—using the ``terminal_observation`` keys in the info dicts returned by the ``VecEnv``.


.. warning::

  When defining a custom ``VecEnv`` (for instance, using gym3 ``ProcgenEnv``), you should provide ``terminal_observation`` keys in the info dicts returned by the ``VecEnv``
  (cf. note above).


.. warning::

    When using ``SubprocVecEnv``, users must wrap the code in an ``if __name__ == "__main__":`` if using the ``forkserver`` or ``spawn`` start method (default on Windows).
    On Linux, the default start method is ``fork`` which is not thread safe and can create deadlocks.

    For more information, see Python's `multiprocessing guidelines <https://docs.python.org/3/library/multiprocessing.html#the-spawn-and-forkserver-start-methods>`_.


Vectorized Environments Wrappers
--------------------------------

If you want to alter or augment a ``VecEnv`` without redefining it completely (e.g. stack multiple frames, monitor the ``VecEnv``, normalize the observation, ...), you can use ``VecEnvWrapper`` for that.
They are the vectorized equivalents (i.e., they act on multiple environments at the same time) of ``gym.Wrapper``.

You can find below an example for extracting one key from the observation:

.. code-block:: python

	import numpy as np

	from stable_baselines3.common.vec_env.base_vec_env import VecEnv, VecEnvStepReturn, VecEnvWrapper


	class VecExtractDictObs(VecEnvWrapper):
	    """
	    A vectorized wrapper for filtering a specific key from dictionary observations.
	    Similar to Gym's FilterObservation wrapper:
	        https://github.com/openai/gym/blob/master/gym/wrappers/filter_observation.py

	    :param venv: The vectorized environment
	    :param key: The key of the dictionary observation
	    """

	    def __init__(self, venv: VecEnv, key: str):
	        self.key = key
	        super().__init__(venv=venv, observation_space=venv.observation_space.spaces[self.key])

	    def reset(self) -> np.ndarray:
	        obs = self.venv.reset()
	        return obs[self.key]

	    def step_async(self, actions: np.ndarray) -> None:
	        self.venv.step_async(actions)

	    def step_wait(self) -> VecEnvStepReturn:
	        obs, reward, done, info = self.venv.step_wait()
	        return obs[self.key], reward, done, info

	env = DummyVecEnv([lambda: gym.make("FetchReach-v1")])
	# Wrap the VecEnv
	env = VecExtractDictObs(env, key="observation")


VecEnv
------

.. autoclass:: VecEnv
  :members:

DummyVecEnv
-----------

.. autoclass:: DummyVecEnv
  :members:

SubprocVecEnv
-------------

.. autoclass:: SubprocVecEnv
  :members:

Wrappers
--------

VecFrameStack
~~~~~~~~~~~~~

.. autoclass:: VecFrameStack
  :members:

StackedObservations
~~~~~~~~~~~~~~~~~~~

.. autoclass:: stable_baselines3.common.vec_env.stacked_observations.StackedObservations
  :members:

StackedDictObservations
~~~~~~~~~~~~~~~~~~~~~~~

.. autoclass:: stable_baselines3.common.vec_env.stacked_observations.StackedDictObservations
  :members:

VecNormalize
~~~~~~~~~~~~

.. autoclass:: VecNormalize
  :members:


VecVideoRecorder
~~~~~~~~~~~~~~~~

.. autoclass:: VecVideoRecorder
  :members:


VecCheckNan
~~~~~~~~~~~~~~~~

.. autoclass:: VecCheckNan
  :members:


VecTransposeImage
~~~~~~~~~~~~~~~~~

.. autoclass:: VecTransposeImage
  :members:

VecMonitor
~~~~~~~~~~~~~~~~~

.. autoclass:: VecMonitor
  :members:

VecExtractDictObs
~~~~~~~~~~~~~~~~~

.. autoclass:: VecExtractDictObs
  :members:
Add doc 2019-09-26 09:46:40 +00:00			`.. _vec_env:`

Rename to stable-baselines3 2020-05-05 13:02:35 +00:00			`.. automodule:: stable_baselines3.common.vec_env`
Add doc 2019-09-26 09:46:40 +00:00
			`Vectorized Environments`
			`=======================`

			`Vectorized Environments are a method for stacking multiple independent environments into a single environment.`
Add base doc 2020-05-07 08:10:51 +00:00			Instead of training an RL agent on 1 environment per step, it allows us to train it on ``n`` environments per step.
			Because of this, ``actions`` passed to the environment are now a vector (of dimension ``n``).
			It is the same for ``observations``, ``rewards`` and end of episode signals (``dones``).
			In the case of non-array observation spaces such as ``Dict`` or ``Tuple``, where different sub-spaces
			may have different shapes, the sub-observations are vectors (of dimension ``n``).
Add doc 2019-09-26 09:46:40 +00:00
			`============= ======= ============ ======== ========= ================`
			Name ``Box`` ``Discrete`` ``Dict`` ``Tuple`` Multi Processing
			`============= ======= ============ ======== ========= ================`
			`DummyVecEnv ✔️ ✔️ ✔️ ✔️ ❌️`
			`SubprocVecEnv ✔️ ✔️ ✔️ ✔️ ✔️`
			`============= ======= ============ ======== ========= ================`

			`.. note::`

			`Vectorized environments are required when using wrappers for frame-stacking or normalization.`

			`.. note::`

			`When using vectorized environments, the environments are automatically reset at the end of each episode.`
			Thus, the observation returned for the i-th environment when ``done[i]`` is true will in fact be the first observation of the next episode, not the last observation of the episode that has just terminated.
Support for `VecMonitor` for gym3-style environments (#311) * add vectorized monitor * auto format of the code * add documentation and VecExtractDictObs * refactor and add test cases * add test cases and format * avoid circular import and fix doc * fix type * fix type * oops * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * add test cases * update changelog * fix mutable argument * quick fix * Apply suggestions from code review * fix terminal observation for gym3 envs * delete comment * Update doc and bump version * Add warning when already using `Monitor` wrapper * Update vecmonitor tests * Fixes Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> 2021-04-13 16:09:31 +00:00			You can access the "real" final observation of the terminated episode—that is, the one that accompanied the ``done`` event provided by the underlying environment—using the ``terminal_observation`` keys in the info dicts returned by the ``VecEnv``.

Add doc 2019-09-26 09:46:40 +00:00
			`.. warning::`

Support for `VecMonitor` for gym3-style environments (#311) * add vectorized monitor * auto format of the code * add documentation and VecExtractDictObs * refactor and add test cases * add test cases and format * avoid circular import and fix doc * fix type * fix type * oops * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * add test cases * update changelog * fix mutable argument * quick fix * Apply suggestions from code review * fix terminal observation for gym3 envs * delete comment * Update doc and bump version * Add warning when already using `Monitor` wrapper * Update vecmonitor tests * Fixes Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> 2021-04-13 16:09:31 +00:00			When defining a custom ``VecEnv`` (for instance, using gym3 ``ProcgenEnv``), you should provide ``terminal_observation`` keys in the info dicts returned by the ``VecEnv``
			`(cf. note above).`


			`.. warning::`

			When using ``SubprocVecEnv``, users must wrap the code in an ``if __name__ == "__main__":`` if using the ``forkserver`` or ``spawn`` start method (default on Windows).
			On Linux, the default start method is ``fork`` which is not thread safe and can create deadlocks.

			For more information, see Python's `multiprocessing guidelines <https://docs.python.org/3/library/multiprocessing.html#the-spawn-and-forkserver-start-methods>`_.
Add doc 2019-09-26 09:46:40 +00:00

Documentation fixes (#514) * Update multiprocessing example * Add VecEnvWrapper example * Update docs/guide/vec_envs.rst Co-authored-by: Anssi <kaneran21@hotmail.com> Co-authored-by: Anssi <kaneran21@hotmail.com> 2021-07-18 18:51:41 +00:00			`Vectorized Environments Wrappers`
			`--------------------------------`

			If you want to alter or augment a ``VecEnv`` without redefining it completely (e.g. stack multiple frames, monitor the ``VecEnv``, normalize the observation, ...), you can use ``VecEnvWrapper`` for that.
			They are the vectorized equivalents (i.e., they act on multiple environments at the same time) of ``gym.Wrapper``.

			`You can find below an example for extracting one key from the observation:`

			`.. code-block:: python`

			`import numpy as np`

			`from stable_baselines3.common.vec_env.base_vec_env import VecEnv, VecEnvStepReturn, VecEnvWrapper`


			`class VecExtractDictObs(VecEnvWrapper):`
			`"""`
			`A vectorized wrapper for filtering a specific key from dictionary observations.`
			`Similar to Gym's FilterObservation wrapper:`
			`https://github.com/openai/gym/blob/master/gym/wrappers/filter_observation.py`

			`:param venv: The vectorized environment`
			`:param key: The key of the dictionary observation`
			`"""`

			`def __init__(self, venv: VecEnv, key: str):`
			`self.key = key`
			`super().__init__(venv=venv, observation_space=venv.observation_space.spaces[self.key])`

			`def reset(self) -> np.ndarray:`
			`obs = self.venv.reset()`
			`return obs[self.key]`

			`def step_async(self, actions: np.ndarray) -> None:`
			`self.venv.step_async(actions)`

			`def step_wait(self) -> VecEnvStepReturn:`
			`obs, reward, done, info = self.venv.step_wait()`
			`return obs[self.key], reward, done, info`

			`env = DummyVecEnv([lambda: gym.make("FetchReach-v1")])`
			`# Wrap the VecEnv`
			`env = VecExtractDictObs(env, key="observation")`


Add doc 2019-09-26 09:46:40 +00:00			`VecEnv`
			`------`

			`.. autoclass:: VecEnv`
			`:members:`

			`DummyVecEnv`
			`-----------`

			`.. autoclass:: DummyVecEnv`
			`:members:`

			`SubprocVecEnv`
			`-------------`

			`.. autoclass:: SubprocVecEnv`
			`:members:`

			`Wrappers`
			`--------`

			`VecFrameStack`
			`~~~~~~~~~~~~~`

			`.. autoclass:: VecFrameStack`
			`:members:`

Dictionary Observations (#243) * First commit * Fixing missing refs from a quick merge from master * Reformat * Adding DictBuffers * Reformat * Minor reformat * added slow dict test. Added SACMultiInputPolicy for future. Added private static image transpose helper to common policy * Ran black on buffers * Ran isort * Adding StackedObservations classes used within VecStackEnvs wrappers. Made test_dict_env shorter and removed slow * Running isort :facepalm * Fixed typing issues * Adding docstrings and typing. Using util for moving data to device. * Fixed trailing commas * Fix types * Minor edits * Avoid duplicating code * Fix calls to parents * Adding assert to buffers. Updating changelong * Running format on buffers * Adding multi-input policies to dqn,td3,a2c. Fixing warnings. Fixed bug with DictReplayBuffer as Replay buffers use only 1 env * Fixing warnings, splitting is_vectorized_observation into multiple functions based on space type * Created envs folder in common. Updated imports. Moved stacked_obs to vec_env folder * Moved envs to envs directory. Moved stacked obs to vec_envs. Started update on documentation * Fixes * Running code style * Update docstrings on torch_layers * Decapitalize non-constant variables * Using NatureCNN architecture in combined extractor. Increasing img size in multi input env. Adding memory reduction in test * Update doc * Update doc * Fix format * Removing NineRoom env. Using nested preprocess. Removing mutable default args * running code style * Passing channel check through to stacked dict observations. * Running black * Adding channel control to SimpleMultiObsEnv. Passing check_channels to CombinedExtractor * Remove optimize memory for dict buffers * Update doc * Move identity env * Minor edits + bump version * Update doc * Fix doc build * Bug fixes + add support for more type of dict env * Fixes + add multi env test * Add support for vectranspose * Fix stacked obs for dict and add tests * Add check for nested spaces. Fix dict-subprocvecenv test * Fix (single) pytype error * Simplify CombinedExtractor * Fix tests * Fix check * Merge branch 'master' into feat/dict_observations * Fix for net_arch with dict and vector obs * Fixes * Add consistency test * Update env checker * Add some docs on dict obs * Update default CNN feature vector size * Refactor HER (#351) * Start refactoring HER * Fixes * Additional fixes * Faster tests * WIP: HER as a custom replay buffer * New replay only version (working with DQN) * Add support for all off-policy algorithms * Fix saving/loading * Remove ObsDictWrapper and add VecNormalize tests with dict * Stable-Baselines3 v1.0 (#354) * Bump version and update doc * Fix name * Apply suggestions from code review Co-authored-by: Adam Gleave <adam@gleave.me> * Update docs/index.rst Co-authored-by: Adam Gleave <adam@gleave.me> * Update wording for RL zoo Co-authored-by: Adam Gleave <adam@gleave.me> * Add gym-pybullet-drones project (#358) * Update projects.rst Added gym-pybullet-drones * Update projects.rst Longer title underline * Update changelog Co-authored-by: Antonin Raffin <antonin.raffin@ensta.org> * Include SuperSuit in projects (#359) * include supersuit * longer title underline * Update changelog.rst * Fix default arguments + add bugbear (#363) * Fix potential bug + add bug bear * Remove unused variables * Minor: version bump * Add code of conduct + update doc (#373) * Add code of conduct * Fix DQN doc example * Update doc (channel-last/first) * Apply suggestions from code review Co-authored-by: Anssi <kaneran21@hotmail.com> * Apply suggestions from code review Co-authored-by: Adam Gleave <adam@gleave.me> Co-authored-by: Anssi <kaneran21@hotmail.com> Co-authored-by: Adam Gleave <adam@gleave.me> * Make installation command compatible with ZSH (#376) * Add quotes * Add Zsh bracket info * Add clarify pip installation line * Make note bold * Add Zsh pip installation note * Add handle timeouts param * Fixes * Fixes (buffer size, extend test) * Fix `max_episode_length` redefinition * Fix potential issue * Add some docs on dict obs * Fix performance bug * Fix slowdown * Add package to install (#378) * Add package to install * Update docs packages installation command Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Fix backward compat + add test * Fix VecEnv detection * Update doc * Fix vec env check * Support for `VecMonitor` for gym3-style environments (#311) * add vectorized monitor * auto format of the code * add documentation and VecExtractDictObs * refactor and add test cases * add test cases and format * avoid circular import and fix doc * fix type * fix type * oops * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * add test cases * update changelog * fix mutable argument * quick fix * Apply suggestions from code review * fix terminal observation for gym3 envs * delete comment * Update doc and bump version * Add warning when already using `Monitor` wrapper * Update vecmonitor tests * Fixes Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Reformat * Fixed loading of ``ent_coef`` for ``SAC`` and ``TQC``, it was not optimized anymore (#392) * Fix ent coef loading bug * Add test * Add comment * Reuse save path * Add test for GAE + rename `RolloutBuffer.dones` for clarification (#375) * Fix return computation + add test for GAE * Rename `last_dones` to `episode_starts` for clarification * Revert advantage * Cleanup test * Rename variable * Clarify return computation * Clarify docs * Add multi-episode rollout test * Reformat Co-authored-by: Anssi "Miffyli" Kanervisto <kaneran21@hotmail.com> * Fixed saving of `A2C` and `PPO` policy when using gSDE (#401) * Improve doc and replay buffer loading * Add support for images * Fix doc * Update Procgen doc * Update changelog * Update docstrings Co-authored-by: Adam Gleave <adam@gleave.me> Co-authored-by: Jacopo Panerati <jacopo.panerati@utoronto.ca> Co-authored-by: Justin Terry <justinkterry@gmail.com> Co-authored-by: Anssi <kaneran21@hotmail.com> Co-authored-by: Tom Dörr <tomdoerr96@gmail.com> Co-authored-by: Tom Dörr <tom.doerr@tum.de> Co-authored-by: Costa Huang <costa.huang@outlook.com> * Update doc and minor fixes * Update doc * Added note about MultiInputPolicy in error of NatureCNN * Merge branch 'master' into feat/dict_observations * Address comments * Naming clarifications * Actually saving the file would be nice * Fix edge case when doing online sampling with HER * Cleanup * Add sanity check Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> Co-authored-by: Anssi "Miffyli" Kanervisto <kaneran21@hotmail.com> Co-authored-by: Adam Gleave <adam@gleave.me> Co-authored-by: Jacopo Panerati <jacopo.panerati@utoronto.ca> Co-authored-by: Justin Terry <justinkterry@gmail.com> Co-authored-by: Tom Dörr <tomdoerr96@gmail.com> Co-authored-by: Tom Dörr <tom.doerr@tum.de> Co-authored-by: Costa Huang <costa.huang@outlook.com> 2021-05-11 10:29:30 +00:00			`StackedObservations`
			`~~~~~~~~~~~~~~~~~~~`

			`.. autoclass:: stable_baselines3.common.vec_env.stacked_observations.StackedObservations`
			`:members:`

			`StackedDictObservations`
			`~~~~~~~~~~~~~~~~~~~~~~~`

			`.. autoclass:: stable_baselines3.common.vec_env.stacked_observations.StackedDictObservations`
			`:members:`
Add doc 2019-09-26 09:46:40 +00:00
			`VecNormalize`
			`~~~~~~~~~~~~`

			`.. autoclass:: VecNormalize`
			`:members:`
Add base doc 2020-05-07 08:10:51 +00:00

			`VecVideoRecorder`
			`~~~~~~~~~~~~~~~~`

			`.. autoclass:: VecVideoRecorder`
			`:members:`


			`VecCheckNan`
			`~~~~~~~~~~~~~~~~`

			`.. autoclass:: VecCheckNan`
			`:members:`


			`VecTransposeImage`
			`~~~~~~~~~~~~~~~~~`

			`.. autoclass:: VecTransposeImage`
			`:members:`
Support for `VecMonitor` for gym3-style environments (#311) * add vectorized monitor * auto format of the code * add documentation and VecExtractDictObs * refactor and add test cases * add test cases and format * avoid circular import and fix doc * fix type * fix type * oops * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * Update stable_baselines3/common/monitor.py Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> * add test cases * update changelog * fix mutable argument * quick fix * Apply suggestions from code review * fix terminal observation for gym3 envs * delete comment * Update doc and bump version * Add warning when already using `Monitor` wrapper * Update vecmonitor tests * Fixes Co-authored-by: Antonin RAFFIN <antonin.raffin@ensta.org> 2021-04-13 16:09:31 +00:00
			`VecMonitor`
			`~~~~~~~~~~~~~~~~~`

			`.. autoclass:: VecMonitor`
			`:members:`

			`VecExtractDictObs`
			`~~~~~~~~~~~~~~~~~`

			`.. autoclass:: VecExtractDictObs`
			`:members:`