[论文] Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Iden...
研究领域: ML 作者: Yisen Xi 发布时间: 2025-09-01 arXiv: 2509.00143
论文概要
研究领域: ML 作者: Yisen Xi 发布时间: 2025-09-01 arXiv: 2509.00143
中文摘要
2025-2026年的AI市场出现了一波隐身发布浪潮:前沿模型以代号匿名发布在开发者平台上。对于用户而言,身份决定了数据处理条款、供应链风险和能力预期。目前尚无经验证的方法论用于匿名模型的黑盒身份验证:从业者检查清单缺乏准确性证据,而自我标识在设计上就不可信。我们提出了一种针对API服务模型的四阶段取证审计协议。第0阶段从归档平台快照(互联网档案馆)重建发布时配置,暴露预览-生产漂移。第1阶段针对平台目录进行配置指纹采集(上下文、输出上限、推理、模态)。第2阶段使用跨长度差异测试分词器身份,拒绝短提示冲突。第3阶段用行为探针进行佐证。我们在10个已知身份发布上测试声明一致性(7个精确、2个精度差异、1个部分、0个反向),而非匿名状态下的端到端识别。前瞻性验证在一个旗舰案例上:其2026-08-23分析指向GLM-5.3版本线,官方揭示确认了这些家族和版本线推断(部署变体未预先断言;Flash在揭示后一致),以及三个仅第0阶段的案例,其中协议产生分级假设或拒绝而非猜测。提供了仅标准库的实现作为补充材料。
原文摘要
The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No validated methodology exists for black-box identity verification of anonymous models: practitioner checklists lack accuracy evidence, and self-identification is untrustworthy by design. We propose a four-stage forensic audit protocol for API-served models. Stage 0 reconstructs launch-time configuration from archived platform snapshots (Internet Archive), exposing preview--production drift. Stage 1 fingerprints configuration (context, output ceiling, reasoning, modality) against the platform catalog. Stage 2 tests tokenizer identity with a cross-l...
*自动采集于 2026-09-02*
#论文 #arXiv #ML #小凯