返回模型
说明文档
This model serves as the baseline for the Ocean Plastic Collection environment, trained and tested on task <code>2</code> using the Proximal Policy Optimization (PPO) algorithm.<br> <br> Environment: Ocean Plastic Collection<br> Task: <code>2</code><br> Algorithm: <code>PPO</code><br> Episode Length: <code>5000</code><br> Training <code>max_steps</code>: <code>3000000</code><br> Testing <code>max_steps</code>: <code>150000</code><br> <br> Train & Test Scripts<br> Download the Environment
hivex-research/hivex-OPC-PPO-baseline-task-2
作者 hivex-research
reinforcement-learning
hivex
↓ 0
♥ 0
创建时间: 2024-08-28 21:57:11+00:00
更新时间: 2025-03-20 23:06:45+00:00
在 Hugging Face 上查看文件 (18)
.gitattributes
Agent.onnx
ONNX
Agent/Agent-1499962.onnx
ONNX
Agent/Agent-1499962.pt
Agent/Agent-1999995.onnx
ONNX
Agent/Agent-1999995.pt
Agent/Agent-2499935.onnx
ONNX
Agent/Agent-2499935.pt
Agent/Agent-2999856.onnx
ONNX
Agent/Agent-2999856.pt
Agent/Agent-3000012.onnx
ONNX
Agent/Agent-3000012.pt
Agent/checkpoint.pt
Agent/events.out.tfevents.1716785041.RICHARD.39584.0
README.md
configuration.yaml
run_logs/timers.json
run_logs/training_status.json