强化学习 Q-Learning 入门
Q-Learning 通过 Q 表记录状态动作的价值。Agent 探索环境根据奖励更新 Q 值。$\epsilon$-greedy 平衡探索和利用。本文用走迷宫的例子实现 Q-Learning。
16
0
0
2026-07-09