CAICT Agentic AI Capability Assessment Standard 1.0: China Starts Issuing 'Licenses' for Models That Can Actually Work
On July 1, 2026, the China Academy of Information and Communications Technology (CAICT), jointly with the China AI Industry Alliance (AIIA), officially released the *Agentic AI Capability Assessment Standard 1.0* and launched an official evaluation platform. This is the first "Agentic AI system-level evaluation framework" published by a national-level official research institution anywhere in the world.
Unlike LLM benchmarks of the past two years, the evaluation target is not "a model" but "a complete agent system," covering six core capabilities:
- Task planning: decomposing a high-level goal (e.g., "build me a quant backtesting system") into 5–15 executable subtasks, including dependencies, resource invocation, and failure fallback strategies.
- Tool calling: accurate tool invocation, handling tool exceptions, and adjusting next actions based on tool feedback.
- Long-term memory: maintaining context over 100+ turns of multi-turn dialogue with traceable key decisions.
- Environment awareness: understanding real-world state changes (file creation, database writes, API status codes) and adapting behavior accordingly.
- Security and compliance: proactively avoiding unauthorized operations, data leakage, and prompt injection attacks.
- Multi-agent collaboration: correct division of labor, conflict avoidance, and result merging in multi-agent systems.
- Static benchmarking. Real agent performance depends on task complexity, environment noise, and user behavior — a 5-star static rating may be 3 stars in production. This is a fundamental limitation of all benchmarks.
- Policy vs. technical neutrality. As a "national team," CAICT's past standards (e.g., DSMM) were later found to have implicit preferences for domestic solutions. The 1.0 standard shows no visible bias yet, but actual evaluation cases over the next 6–12 months should be watched.
- Unclear international alignment. Interoperability with ISO/IEC JTC1 SC42, IEEE P7000 series, and ETSI is undefined; without mutual recognition, Chinese AI products going overseas face duplicate certification and doubled compliance costs.
Each capability has 3–5 sub-indicators with explicit scoring scales. Every agent system receives a 1–5 star "Agentic AI capability level" — 5 stars means it can operate autonomously in high-risk scenarios like finance, healthcare, and government; 1 star means it can only handle simple single-turn, single-tool tasks.
Why It Matters
This is the dividing line between "chat-capable LLMs" and "work-capable LLMs." Past evaluations (MMLU, HumanEval, SWE-bench, GPQA Diamond) measured how much a model *knows*, not whether it can *finish a task*. A model scoring 70% on SWE-bench may still fail to independently complete a full engineering task in a real environment. CAICT's 1.0 standard shifts the evaluation focus from "model intelligence" to "task completion" — consistent in spirit with the Stanford/Apple paper (July 2) questioning the marginal utility of multi-agent collaboration.
Why China First?
1. China's application layer runs fastest. China has the world's highest AI application penetration in industrial internet, government, healthcare, and legal domains. In 2025, China's ToB AI market reached RMB 450 billion (~60% of the US market), but China's ToB AI skews toward "agent execution / process automation" (>55%) versus the US's chatbot/content generation focus (>60%) — making system-level agentic evaluation a more rigid need. 2. "Standards-first" regulation. The EU's AI Act Phase 2 implementing rules (June 2026) and NIST's AI Risk Management Framework (2025) focus on risk management — telling you what *can't* use AI. CAICT's standard directly assigns capability levels, effectively a proto "market entry license" for agentic AI products.
Impact at Three Levels
1. Forcing engineering direction for Chinese LLM vendors. B2B products from Zhipu, Alibaba, Moonshot, ByteDance, and DeepSeek will need the certification in the next 12 months — turning "can this model work?" from market self-assessment into official certification, with customers ordering directly by star rating. 2. Direct impact on AI coding tools. Overseas tools (Cursor, Claude Code, GitHub Copilot) targeting government/SOE/financial customers in China will almost certainly need certification — but likely won't reach 5 stars due to weak adaptation to local toolchains (DingTalk, Feishu, WeCom, Alibaba Cloud, Huawei Cloud), strengthening domestic tools (Tongyi Lingma, Trae, CodeGeeX, CodeArts). 3. Early regulation of embodied agents. CAICT stated that version 2.0 will add an embodied intelligence evaluation module, meaning humanoid robot makers (Unitree, Zhiyuan, Robot Era) will enter this system before end of 2026, curbing self-reported leaderboard chaos.
Risks
Source: CAICT, *Agentic AI Capability Assessment Standard 1.0 Officially Released*, 2026-07-01, https://www.caict.ac.cn/kxyj/qwfb/bps/202607/t20260701_647283.htm