• •    

基于大语言模型的植物性状自动提取框架构建与评价体系

陈佳乐, 赵颖菲, 昌海超, 董必成, 于飞海   

  1. 北京林业大学, 北京 100083 中国
    北京林业大学,黄河流域生态保护国家林草局重点实验室, 100083
    绍兴大学生命与环境科学学院, 312000
  • 收稿日期:2026-04-03 修回日期:2026-06-04
  • 基金资助:
    国家自然科学基金项目(31500331)

A framework and evaluation system for plant trait extraction based on large language models

Jia-Le Chen, Ying-Fei Zhao, Hai-Chao Chang, Bi-Cheng Dong, Fei-Hai Yu   

  1. , Beijing Forestry University 100083, China
    , The Key Laboratory of Ecological Protection in the Yellow River Basin of National Forestry and Grassland Administration, Beijing Forestry University 100083,
    , School of Life and Environmental Sciences, Shaoxing University 312000,
  • Received:2026-04-03 Revised:2026-06-04

摘要: 【目的】本研究旨在构建一套基于大语言模型(LLMs)的植物性状自动提取方法框架,并系统评估模型类型、提示词设计及参数设置对任务性能的影响。【方法】以《中国植物志》为数据来源,构建包含27,922条人工标注记录的标准化评测数据集,设计统一的数据预处理流程与系统化提示词模板。在此基础上,选取DeepSeek-V3与DeepSeek-R1-0528-Qwen3-8B两类模型,结合不同温度参数及是否使用系统化提示词构建三因素实验设计,在植物生长型与生命周期识别任务中进行多轮重复实验。采用缺失率、幻觉率、准确率、精确率、召回率、特异性、F1分数及AUC值进行综合评价,并通过方差分析检验各因素及其交互作用的显著性。【主要结果】研究表明,不同模型在参数敏感性与错误类型控制上存在差异,其中DeepSeek-V3整体表现更稳定,DeepSeek-R1-0528-Qwen3-8B则对系统化提示词更加敏感;低温度设置总体有助于降低幻觉率和输出不稳定风险,而较高温度(0.6)在部分场景下可能提高召回能力;系统化提示词显著降低模型输出的缺失率和幻觉率,提高了输出完整性与类别合法性,并在部分任务和指标上改善了分类性能。相较于单一模型性能比较,参数配置与提示策略对大语言模型在植物学文本解析任务中的实际可用性具有更关键影响。本研究构建的方法框架可为生物多样性信息自动化提取提供可复用的技术路径。

关键词: 大语言模型, 植物功能性状, 文本挖掘, 生物多样性信息学, 评测体系构建

Abstract: Abstract Aim This study aims to develop a methodological framework for plant trait extraction using large language models (LLMs) and systematically evaluate the effects of model type, prompt design, and parameter settings on task performance. Methods Using the Flora of China as the data source, we constructed a standardized benchmark dataset containing 27,922 manually annotated records and designed a unified data preprocessing pipeline and systematic prompt templates. Two models (DeepSeek-V3 and DeepSeek-R1-0528-Qwen3-8B) were selected, and a three-factor experimental design was established by combining model type, temperature setting, and prompt strategy. Multiple repeated experiments were conducted for plant growth form and life span classification tasks. Model performance was comprehensively evaluated using missing rate, mismatch rate, accuracy, precision, recall, specificity, F1-score, and AUC, and analysis of variance was used to test the significance of main effects and interaction effects. Important findings The results showed that different models varied in parameter sensitivity and error-type control. DeepSeek-V3 showed more stable overall performance, whereas DeepSeek-R1-0528-Qwen3-8B was more sensitive to systematic prompts. A lower temperature setting generally helped reduce hallucination rates and the risk of output instability, whereas a higher temperature setting (0.6) may improve recall in some scenarios. Systematic prompts significantly reduced missing and hallucination rates, improved output completeness and category validity, and enhanced classification performance in some tasks and metrics. Compared with simple model performance comparison, parameter configuration and prompt strategy had a more critical influence on the practical applicability of LLMs in botanical text parsing tasks. The proposed framework provides a reusable technical pathway for automated biodiversity information extraction.

Key words: large language models, plant functional traits, text mining, biodiversity informatics, evaluation framework construction