论文概要
研究领域: ML
作者: Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan, Mohammad Alizadeh, Fred Kjolstad, Suvinay Subramanian
发布时间: 2026-09-04
arXiv: 2609.05364
中文摘要
机器学习性能建模对长寿软件是独特敌意的领域:今天抽象中烘焙的假设被明天的模型和系统推翻,迫使性能建模框架不断重构。同时,AI编码智能体已足够快速和有能力,以至于重新生成整个库比渐进修补的技术债务更便宜。本文描述SMART,一个ML系统的严格符号性能建模库,其主分支几乎不含代码:仓库是自包含自然语言设计文档的DAG,编码子智能体仅从新版本更新的文档重新生成实现,每个人类变更都是对文档的自然语言编辑——自文档化。两个要素使重新生成可靠:(i)围绕逐步工作示例构建的设计文档风格,作为生成智能体的上下文演示;(ii)最小递归定义的操作IR,具有符号(SymPy)成本表达式、用于大规模扫描的快速分析汇总模式和用于细粒度调度研究的慢速模调度模式。重新生成的实现复现手工审计参考模型——包括TPU pod切片上的DeepSeek-V3 serving——达到舍入精度,表明设计文档——而非代码——可以成为ML系统协同设计工具的持久产物。
原文摘要
Machine-learning performance modeling is a uniquely hostile terrain for long-lived software: the assumptions baked into today's abstractions are invalidated by tomorrow's models and systems, forcing perpetual refactoring of performance-modeling frameworks. Meanwhile, AI coding agents have become fast and capable enough that regenerating an entire library is cheaper than paying down the tech debt of incrementally patching it. We describe SMART, a rigorous symbolic performance-modeling library for ML systems whose main branch contains almost no code: the repository is a DAG of self-contained natural-language design docs, coding sub-agents regenerate the implementation from only the docs on new version updates, and every human change is a natural-language edit to a doc--self-documenting by co...
自动采集于 2026-09-09
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。