基于世界模型自校准的智能体终身学习方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金面上/青年项目 (62602156, 62476068, 62306092, 62502115); 山东省自然科学基金重大基础研究项目 (ZR2025ZD01); 山东省自然科学基金 (ZR2024QF066, ZR2025QC1516, ZR2025QC1520)


A Lifelong Learning Method for Agents Based on World-model-driven Self-calibration
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    在大规模多模态模型驱动的移动智能体研究中,构建内在世界模型以实现图形用户界面(GUI)环境下长程推理与规划能力的重要基础.然而,开放环境的动态非确定性常导致智能体产生状态转移幻觉,即内在预测预期与环境真实反馈失配,进而引发规划失效与错误级联.现有的自我反思机制多受限于模型自身的认知盲区,且缺乏在长周期交互中自主演进的能力.鉴于此,本文提出一种基于System-1和System-2解耦的智能体双层认知系统结构.System-1作为直觉执行层,利用参数化世界模型进行快速状态推演与决策树规划;System-2作为监督认知层,通过实时校验状态转移一致性扮演智能体的认知监管层.当检测到动力学残差超过阈值时,System-2介入并利用反事实因果归因机制定位失效根因,引导智能体执行状态回溯与动态重规划.此外,该架构通过构建情景因果记忆库形成了闭环数据飞轮,驱动世界模型参数在交互中持续自校准.在本文构建的多领域跨应用GUI任务基准及跨域迁移设置下,实验结果表明该架构能够降低错误级联并促进部分因果知识内化.随着交互规模的扩大,系统干预频率下降而任务独立成功率提升,显示出在所测试GUI任务范围内进行持续自校准的阶段性效果.源码清单见链接:https://github.com/LuRu520/rs025_materials.git

    Abstract:

    In research on mobile agents driven by large multimodal models, constructing an internal world model constitutes an important foundation for enabling long-horizon reasoning and planning in graphical user interface (GUI) environments. However, the dynamic and nondeterministic nature of open environments often causes agents to produce state-transition hallucinations, namely, mismatches between internal predictions and actual environmental feedback, which can in turn lead to planning failures and cascading errors. Existing self-reflection mechanisms are often constrained by the model's own cognitive blind spots and lack the capacity for autonomous evolution over prolonged interactions. To address these limitations, this paper proposes a dual-layer cognitive architecture for agents based on the decoupling of System-1 and System-2. System-1 serves as the intuitive execution layer and uses a parameterized world model for rapid state rollouts and decision-tree planning; System-2 serves as the supervisory cognitive layer, providing cognitive oversight through real-time verification of state-transition consistency. When the dynamics residual exceeds a threshold, System-2 intervenes and uses a counterfactual causal attribution mechanism to identify the root cause of the failure, guiding the agent to perform state backtracking and dynamic replanning. Furthermore, by constructing an episodic causal memory bank, the architecture creates a closed-loop data flywheel that drives the continual self-calibration of world model parameters during interaction. Experiments on the multi-domain, cross-application GUI task benchmark developed in this work and in a cross-domain transfer setting show that the architecture can reduce error cascades and facilitate the partial internalization of causal knowledge. As the interaction scale increases, the frequency of System-2 interventions decreases while the independent task success rate increases, providing preliminary evidence of continual self-calibration within the scope of the tested GUI tasks.

    参考文献
    相似文献
    引证文献
引用本文

卢小芬,张北辰,郑培泉,戚兆波,刘心岩,张维刚.基于世界模型自校准的智能体终身学习方法.软件学报,2027,38(5):

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-04-27
  • 最后修改日期:2026-08-20
  • 录用日期:
  • 在线发布日期: 2026-09-14
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号