数据工厂的构成、建设模式和运营机制研究OA
Research on the composition,construction models and operation mechanisms of data factories
高质量数据集是人工智能大模型训练的核心"燃料".当前,高质量数据集构建主要由人工智能企业自行完成,呈现零散化、作坊式、非标化的特点,难以满足人工智能大模型快速发展的需求.借鉴水厂、电厂等资源型基础设施的发展规律,结合国内外高质量数据集设施化生产的典型实践,提出"数据工厂"概念,将其定义为面向人工智能大模型应用、设施化规模化构建高质量数据集的生产设施.系统阐述了数据工厂由"储备车间""生产车间""中试车间"构成的三级架构体系,分析了数据标注企业升级、数据存储基地转型、人工智能企业延伸和技术企业创新设立四种建设模式,提出了保障模式、定制模式、电商模式和结对子模式四种运营机制,为推动高质量数据集设施化、规模化供给提供理论支撑和实践参考.
High-quality datasets are the core fuel for training large AI models.Currently,the construction of high-quality datasets is mainly carried out by AI enterprises themselves,which presents the characteristics of fragmentation,workshop-style operation and non-standardization,making it difficult to meet the rapid development needs of large AI models.Drawing on the development patterns of resource-based infrastruc-ture such as water and power plants,and combining domestic and international best practices in facility-based production,this paper proposes the concept of"data factory",defining it as a production facility specifically designed for the application of large AI models and for the facility-based,large-scale construction of high-quality datasets.The paper systematically expounds the three-level architecture system of the data facto-ry,which consists of storage workshop,production workshop,and pilot workshop.Four construction models and four operation mechanisms are proposed,providing theoretical support and practical references for promoting the facility-based and large-scale supply of high-quality datasets.
涂群;耿贵宁;张茜茜
北京化工大学 经济管理学院,北京 100029三六零数字安全科技集团有限公司,北京 100015北京物资学院 计算机与人工智能学院,北京 101126
管理科学
数据工厂高质量数据集数据基础设施数据要素
data factoryhigh-quality datasetdata infrastructuredata element
《网络安全与数据治理》 2026 (4)
9-16,8
北京市社会科学基金(23GLC058)
评论