{
    "created": "2026-07-01 16:48:15",
    "updated": "2026-08-19 17:10:01",
    "id": "e60ccb9a-ecf2-4e0a-963d-02d66926768f",
    "version": 4,
    "ds_topic": null,
    "title_cn": "多无人机协同围捕训练与轨迹仿真结果数据集",
    "title_en": "Multi-UAV Cooperative Encirclement Training and Trajectory Simulation Results Dataset",
    "ds_abstract": "<p>&emsp;&emsp;本数据集面向受限二维空间内多无人机协同围捕多目标任务中的深度强化学习训练过程分析与轨迹仿真展示，数据来源于论文《Deep Reinforcement Learning-Based Cooperative Encirclement Strategy for Multi-UAVs》的仿真实验。论文在集中训练、分散执行（CTDE）框架下，构建了阶段奖励与多任务驱动的MADDPG算法，通过将围捕任务分解为接近、跟踪、包围和捕获等连续阶段，缓解多智能体围捕任务中的奖励稀疏问题；同时结合贪婪算法进行目标分配，并为目标无人机设计基于人工势场思想的逃逸策略，以形成更贴近对抗场景的动态仿真环境。实验基于Python、PyTorch和OpenAI Gym的多智能体仿真环境开展，训练六架围捕无人机对两个目标无人机进行协同围捕，并在动态障碍物条件下评估训练后策略的表现。汇交文件按metadata、data、images目录组织，共包含1份元数据表、2个CSV数据文件、2张PNG可视化图片和1份README说明文件；核心数据实体包括data/score_history.csv、data/score_history_original_row62.csv、images/MADDPG_Training_Reward.png和images/Encirclement_Trajectory.png。训练奖励CSV包含episode、reward和moving_average_100三个字段，其中episode为1至10000的回合编号，reward为单回合累计奖励，moving_average_100为100回合滑动平均奖励；reward和moving_average_100均为仿真奖励值，无物理量纲。该数据集可用于分析多智能体深度强化学习训练收敛过程、阶段奖励与多任务机制的效果，以及多无人机协同围捕策略在复杂动态环境中的轨迹表现。</p>",
    "ds_source": "<p>&emsp;&emsp;本数据集由论文相关仿真实验产生，为完全仿真生成数据，不来源于实测观测、购买、交换、镜像或第三方下载数据。</p>",
    "ds_process_way": "<p>&emsp;&emsp;本数据集基于二维多无人机协同围捕仿真环境和MADDPG训练流程生成。训练阶段以回合为单位运行仿真，记录每回合训练累计奖励，并整理为data/score_history.csv；同时保留从源score_history.csv中提取的原始单行奖励序列data/score_history_original_row62.csv，用于追溯。训练奖励曲线根据10000回合奖励序列进行100回合滑动平均后绘制，输出为images/MADDPG_Training_Reward.png。轨迹结果由训练后策略在包含动态障碍的仿真环境中评估得到，并通过可视化程序输出围捕轨迹图images/Encirclement_Trajectory.png。</p>",
    "ds_quality": "<p>&emsp;&emsp;本数据集在生成过程中进行了完整性和可读性检查。训练奖励数据文件可正常读取，核心奖励曲线对应10000回合训练记录。规范化后的score_history.csv共10000行，字段包括episode、reward、moving_average_100；episode取值范围为1至10000，reward取值范围约为-3823.6044至3134.3994，moving_average_100有效值为9901个，取值范围约为-2336.4726至2.8238；reward和moving_average_100均为仿真奖励值，无物理量纲。两张PNG图像分辨率均为1600×1200，可直接展示训练奖励变化和训练后围捕轨迹结果。数据内容与论文中的训练奖励曲线和围捕轨迹结果相对应，未包含与本数据集无关的地理空间实测信息。</p>",
    "ds_acq_start_time": "2025-01-01 00:00:00",
    "ds_acq_end_time": null,
    "ds_acq_place": "南京",
    "ds_acq_lon_east": null,
    "ds_acq_lat_south": null,
    "ds_acq_lon_west": null,
    "ds_acq_lat_north": null,
    "ds_acq_alt_low": null,
    "ds_acq_alt_high": null,
    "ds_share_type": "open-access",
    "ds_total_size": 874010,
    "ds_files_count": 6,
    "ds_format": "csv, png",
    "ds_space_res": "",
    "ds_time_res": "",
    "ds_coordinate": "无",
    "ds_projection": "",
    "ds_thumbnail": "e60ccb9a-ecf2-4e0a-963d-02d66926768f.png",
    "ds_thumb_from": 0,
    "ds_ref_way": "",
    "paper_ref_way": "",
    "ds_ref_instruction": "使用本数据集时，请注明数据集名称，并建议同时引用相关论文《Deep Reinforcement Learning-Based Cooperative Encirclement Strategy for Multi-UAVs》。数据文件按metadata、data和images目录组织：data目录包含规范化训练奖励数据和原始奖励序列追溯文件，images目录包含训练奖励曲线图和围捕轨迹图。score_history.csv中的episode表示训练回合编号，reward表示单回合累计奖励，moving_average_100表示100回合滑动平均奖励；reward和moving_average_100均无物理量纲。数据主要用于多无人机协同围捕、深度强化学习训练过程分析和仿真轨迹展示，不包含实测地理空间数据。",
    "ds_from_station": "",
    "organization_id": "9ecaaa78-39e9-411e-9f24-274e12aa643f",
    "ds_serv_man": "王和",
    "ds_serv_phone": "15105193069",
    "ds_serv_mail": "wanghe91@seu.edu.cn",
    "doi_value": "",
    "subject_codes": [
        "410"
    ],
    "quality_level": 0,
    "publish_time": "2026-07-09 10:57:54",
    "last_updated": "2026-07-21 09:24:56",
    "protected": false,
    "protected_to": "2028-06-30 00:00:00",
    "lang": "zh",
    "cstr": "11738.11.ncdc.db7477.2026",
    "i18n": {
        "en": {
            "title": "Multi-UAV Cooperative Encirclement Training and Trajectory Simulation Results Dataset",
            "ds_format": "csv, png",
            "ds_source": "This dataset was produced from the simulation experiments associated with the paper. It is fully generated through simulation and does not originate from field observations, purchased data, exchanged data, mirrored datasets, or third-party downloads.",
            "ds_quality": "Completeness and readability checks were performed during dataset preparation. The training reward CSV file is readable, and the core reward curve corresponds to a 10,000-episode training record. The normalized score_history.csv contains 10,000 rows with three fields: episode, reward, and moving_average_100. The episode values range from 1 to 10,000, reward values range approximately from -3823.6044 to 3134.3994, and moving_average_100 has 9,901 valid values ranging approximately from -2336.4726 to 2.8238. Both reward and moving_average_100 are dimensionless simulation reward values. Both PNG images have a resolution of 1600 by 1200 pixels and can directly present the training reward trend and the post-training encirclement trajectory result. The dataset corresponds to the training reward curve and encirclement trajectory result reported in the paper and does not contain unrelated field-based geospatial measurements.",
            "ds_ref_way": "",
            "ds_abstract": "This dataset supports the analysis of deep reinforcement learning training processes and trajectory-simulation visualization for cooperative multi-UAV multi-target encirclement in a constrained two-dimensional space. The data are derived from the simulation experiments reported in the paper \"Deep Reinforcement Learning-Based Cooperative Encirclement Strategy for Multi-UAVs\". Under the centralized-training and decentralized-execution (CTDE) framework, the paper develops a stage-reward and multi-task-driven MADDPG algorithm. The encirclement mission is decomposed into continuous stages such as approaching, tracking, surrounding, and capturing, thereby alleviating reward sparsity in multi-agent encirclement tasks. A greedy algorithm is used for target assignment, and an artificial-potential-field-inspired escape strategy is designed for the target UAVs, forming a dynamic adversarial simulation environment. The experiments were conducted in a Python, PyTorch, and OpenAI Gym multi-agent simulation environment, where six encircling UAVs were trained to cooperatively encircle two target UAVs and the trained policy was evaluated under dynamic-obstacle conditions. The submitted files are organized in metadata, data, and images directories and include one metadata workbook, two CSV data files, two PNG visualization images, and one README file. The core data entities are data/score_history.csv, data/score_history_original_row62.csv, images/MADDPG_Training_Reward.png, and images/Encirclement_Trajectory.png. The training reward CSV contains three fields: episode, reward, and moving_average_100. The episode field records episode numbers from 1 to 10,000, reward records the accumulated reward of each episode, and moving_average_100 records the 100-episode moving-average reward. Both reward and moving_average_100 are dimensionless simulation reward values. This dataset can be used to analyze training convergence in multi-agent deep reinforcement learning, the effects of stage-reward and multi-task mechanisms, and trajectory performance of cooperative multi-UAV encirclement strategies in complex dynamic environments.",
            "ds_time_res": "",
            "ds_acq_place": "Nanjing",
            "ds_space_res": "",
            "ds_projection": "",
            "ds_process_way": "The dataset was generated from a two-dimensional cooperative multi-UAV encirclement simulation environment and the MADDPG training procedure. During training, simulations were executed episode by episode, and the episode-level accumulated training rewards were recorded and organized as data/score_history.csv. The original single-row reward sequence extracted from the source score_history.csv is retained as data/score_history_original_row62.csv for traceability. The training reward curve was plotted from a 10,000-episode reward sequence using a 100-episode moving average and exported as images/MADDPG_Training_Reward.png. The trajectory result was obtained by evaluating the trained policy in a simulation environment with dynamic obstacles and was exported as images/Encirclement_Trajectory.png by the visualization program.",
            "ds_ref_instruction": "When using this dataset, please cite the dataset name and preferably also cite the related paper \"Deep Reinforcement Learning-Based Cooperative Encirclement Strategy for Multi-UAVs\". The files are organized in metadata, data, and images directories. The data directory contains the normalized training reward data and the raw reward-sequence traceability file, while the images directory contains the training reward curve and the encirclement trajectory diagram. In score_history.csv, episode denotes the training episode number, reward denotes the accumulated reward of one episode, and moving_average_100 denotes the 100-episode moving-average reward. Both reward and moving_average_100 are dimensionless. The data are intended for cooperative multi-UAV encirclement research, deep reinforcement learning training-process analysis, and simulation-trajectory visualization, and do not contain field-measured geospatial data."
        }
    },
    "submit_center_id": "ncdc",
    "data_level": 0,
    "recommendation_value": 0,
    "license_type": "https://creativecommons.org/licenses/by/4.0/",
    "doi_reg_from": "reg_local",
    "cstr_reg_from": "reg_local",
    "doi_not_reg_reason": null,
    "cstr_not_reg_reason": null,
    "is_paper_in_submitting": false,
    "belong_to_nieer": false,
    "allow_update_data": false,
    "ds_topic_tags": [
        "多无人机",
        "协同围捕",
        "深度强化学习",
        "MADDPG",
        "训练奖励",
        "轨迹仿真"
    ],
    "ds_subject_tags": [
        "工程与技术科学基础学科"
    ],
    "ds_class_tags": [],
    "ds_locus_tags": [],
    "ds_time_tags": [],
    "ds_contributors": [
        {
            "true_name": "马文瑞",
            "email": "2212307745@qq.com",
            "work_for": "东南大学数学学院",
            "country": "中国"
        }
    ],
    "ds_meta_authors": [
        {
            "true_name": "马文瑞",
            "email": "2212307745@qq.com",
            "work_for": "东南大学数学学院",
            "country": "中国"
        }
    ],
    "ds_managers": [
        {
            "true_name": "王和",
            "email": "wanghe91@seu.edu.cn",
            "work_for": "东南大学",
            "country": "中国"
        }
    ],
    "category": "其他"
}