Breaking the deadly triad with a target network

The deadly triad refers to the instability of a reinforcement learning algorithm when it employs off-policy learning, function approximation, and bootstrapping simultaneously. In this paper, we investigate the target network as a tool for breaking the deadly triad, providing theoretical support for...

সম্পূর্ণ বিবরণ

গ্রন্থ-পঞ্জীর বিবরন
প্রধান লেখক:	Zhang, S, Yao, H, Whiteson, S
বিন্যাস:	Conference item
ভাষা:	English
প্রকাশিত:	PMLR 2021

Breaking the deadly triad with a target network

অনুরূপ উপাদানগুলি