基于大语言模型的长方法分解

doi:10.13328/j.cnki.jos.007329

微信服务号

微信订阅号

2025年5月11日 5:04 星期日

首页 > 过刊浏览>2025年第36卷第6期 >2501-2514. DOI:10.13328/j.cnki.jos.007329

PDF HTML阅读 XML下载导出引用引用提醒

基于大语言模型的长方法分解
DOI:
                        10.13328/j.cnki.jos.007329
                    
CSTR:
                        
                    
作者:
                        徐子懋徐子懋
北京理工大学 计算机学院, 北京 100081
在期刊界中查找
在百度中查找
在本站中查找
姜艳杰姜艳杰
北京大学 计算机学院, 北京 100091
在期刊界中查找
在百度中查找
在本站中查找
张宇霞张宇霞
北京理工大学 计算机学院, 北京 100081
在期刊界中查找
在百度中查找
在本站中查找
刘辉刘辉
北京理工大学 计算机学院, 北京 100081
在期刊界中查找
在百度中查找
在本站中查找

                    
作者单位:
作者简介:
通讯作者:姜艳杰,E-mail:yanjiejiang@pku.edu.cn
中图分类号:TP311
基金项目:国家自然科学基金(62172037, 62232003); 中国博士后科学基金(2023M740078)

Large Language Model-Based Decomposition of Long Methods

Author:

XU Zi-Mao
XU Zi-Mao
School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China
在期刊界中查找
在百度中查找
在本站中查找
JIANG Yan-Jie
JIANG Yan-Jie
School of Computer Science, Peking University, Beijing 100091, China
在期刊界中查找
在百度中查找
在本站中查找
ZHANG Yu-Xia
ZHANG Yu-Xia
School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China
在期刊界中查找
在百度中查找
在本站中查找
LIU Hui
LIU Hui
School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China
在期刊界中查找
在百度中查找
在本站中查找

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

资源附件

文章评论

摘要:

长方法及其他类型的代码坏味阻碍了软件应用程序达到最佳的可读性、可重用性和可维护性. 因此, 人们对长方法的自动检测和分解进行了广泛的研究. 虽然这些方法极大地促进了分解, 但其解决方案往往与最优方案存在很大差异. 为此, 调研公开真实长方法数据集中的可自动化部分, 探讨了长方法的分解情况, 并基于调研结果, 提出了一种基于大语言模型的新方法(称为 Lsplitter), 用于自动分解长方法. 对于给定的长方法, Lsplitter会根据启发式规则和大语言模型将该方法分解为一系列短方法. 然而, 大语言模型经常会拆分出相似的方法, 针对大语言模型的分解结果, Lsplitter利用基于位置的算法, 将物理上连续且高度相似的方法合并成一个较长的方法. 最后对这些候选结果进行排序. 对真实Java项目中的2849个长方法进行了实验, 结果表明, 相较传统结合模块化矩阵的方法, Lsplitter的命中率提升了142%, 相较纯基于大语言模型的方法, 命中率提升了7.6%.

关键词:长方法;重构;大语言模型;分解

Abstract:

Long methods, along with other types of code smells, prevent software applications from reaching their optimal readability, reusability, and maintainability. Consequently, automated detection and decomposition of long methods have been widely studied. Although these approaches have significantly facilitated the decomposition, their solutions often differ significantly from the optimal ones. To address this, the automatable portion of the publicly available dataset containing real-world long methods is investigated. Based on the findings of this investigation, a new method (called Lsplitter) based on large language models (LLMs) is proposed in this study for automatically decomposing long methods. For a given long method, the Lsplitter decomposes the method into a series of shorter methods according to heuristic rules and LLMs. However, LLMs often split out similar methods. In response to the decomposition results of LLMs, Lsplitter utilizes a location-based algorithm to merge physically contiguous and highly similar methods into a longer method. Finally, these candidate results are ranked. Experiments are conducted on 2 849 long methods in real Java projects. The experimental results show that compared with the traditional methods combined with a modularity matrix, the hit rate of Lsplitter is improved by 142%, and compared with the methods purely based on LLMs, the hit rate is improved by 7.6%.

Key words:long method;refactoring;LLM;decomposition

引用本文

徐子懋,姜艳杰,张宇霞,刘辉.基于大语言模型的长方法分解.软件学报,2025,36(6):2501-2514

复制

文章指标

点击次数:
下载次数:
HTML阅读次数:
引用次数:

历史

收稿日期:2024-08-26
最后修改日期:2024-10-14
录用日期:
在线发布日期: 2024-12-10
出版日期:

微信服务号

微信订阅号

引用本文

相关视频

分享

文章指标

历史

文章二维码

微信服务号

微信订阅号

引用本文

相关视频

分享

微信扫一扫：分享

文章指标

历史

文章二维码