This dataset was constructed in June 2025 using a Python simulation platform and the MPE (Multi-Agent Particle Environment) navigation task framework, with automated data collection facilitated by Wandb. It is designed to compare the performance differences between two credit assignment mechanisms—Shapley Dividend and the conventional Shapley Value—when embedded into multi-agent reinforcement learning algorithms (SD3PG and SQDDPG). All experiments were conducted under identical network architectures and hyperparameter configurations. For each set of conditions, data were collected across five distinct random seeds, recording the per-episode average reward and collision count during training. Summary statistics, including maximum, minimum, and mean values, were computed to comprehensively reflect the expected performance, optimal bounds, and robustness of the algorithms. The dataset comprises two CSV table files, namely mean_end_reward.csv and mean_ep_collisions.csv, along with additional visualization and derivative metric files, such as training reward curves, collision trend plots, and convergence speed analysis tables. By comparing convergence speed, peak final rewards, and fluctuation ranges across training curves, it can be verified that Shapley Dividend, relative to Shapley Value, more effectively guides agents to form cooperative strategies in navigation tasks. It not only accelerates the evolution of cooperative behavior and reduces collision frequency, but also achieves superior and more stable final collaborative performance. This dataset provides reusable benchmark data and analytical foundations for research on credit assignment mechanisms in multi-agent systems.
| collect time | 2025/01/01 - |
|---|---|
| collect place | Nanjing |
| data size | 57.6 KiB |
| data format | .csv |
This dataset is completely synthetically generated through simulation and does not originate from specific literature, field measurements, or third-party downloads.
This dataset was automatically collected via Wandb using a Python simulation platform. Each set of data comprises statistical summaries from training runs under five different random seeds, including maximum, minimum, and mean values.
This dataset was automatically collected by the system, ensuring completeness and analytical viability. By analyzing the training curves, we can verify the following conclusion: compared with the Shapley value, the Shapley dividend more effectively accelerates the cooperative evolution process and achieves superior cooperative performance.
This work is licensed under
CC BY 4.0 (Creative Commons Attribution 4.0 International License).
| # | title | file size |
|---|---|---|
| 1 | mean_end_reward.csv | 33.1 KiB |
| 2 | mean_ep_collisions.csv | 24.4 KiB |
| # | category | title | author | year |
|---|---|---|---|---|
| 1 | paper | 2026 |
-
-
©Copyright 2005-. Northwest Institute of Eco-Environment and Resources, CAS.
Donggang West Road 320, Lanzhou, Gansu, China (730000)

