Event Overview
On June 7, at the CCF YOCSEF Hangzhou technical forum (supported by Alibaba's ATH-Qwen business group), Zhu Da, head of Qwen's C-end MOS Lab, gave a talk titled "Qwen C-end Agent Harness: Reflections and Practice."
Zhu Da positioned his team as "pragmatists" — not chasing the theoretical ceiling of model capability, but solving the concrete problem of "how to use Qwen models to serve 300+ million monthly active users." The main subject was the "general complex-task Agent" (the "capsule" entry above the Qwen App's chat box), which officially launched in January 2026. He condensed the Agent's engineering methodology into four words: more, faster, better, cheaper (duo, kuai, hao, sheng).
The reported numbers are striking: execution time reduced to 1/3 of the initial version, token consumption at 1/10 or less of comparable overseas products. On a national-scale app with over 100 million MAU, this is a credible engineering scorecard.
Deep Dive
1. "More, faster, better, cheaper" treats the C-end Agent like e-commerce engineering.
- More: supports many task types — information gathering, research and analysis, daily-life services, office collaboration, code development. One general Agent covers all types rather than N vertical workflows.
- Faster: execution time should map to delivery quality — 5 minutes / 10 minutes / 30 minutes of execution should yield stepwise quality gains. Key techniques: task parallelization plus "path solidification" (a first-encountered task requires extensive reflection; after accumulating experience, the execution path is solidified into code).
- Better: the industry's hardest problem isn't "how to do well" but "how to define well." For a task like collecting all coin types in China from the Qin to Qing dynasties, there's no completion standard; ordinary agents stop after finding a dozen. Search-paradigm optimization and context management are needed to "push" the model to continue.
- Cheaper: the product is free, so cost control is the lifeline. Reduce model calls, have the model write code instead of repeated reasoning, solidify paths, optimize caching.
2. "Proactive service" is the most notable C-end Agent paradigm upgrade for H2 2026.
Zhu Da sketched a true personal-assistant profile: not "I always go ask my assistant," but an assistant that knows everything about you and proactively serves and pushes.
The landing architecture has four components: User Memory + Environment (situational awareness) + Task System + Assistant (unified reach).
The biggest challenge is "emotional intelligence." His example: a user previously mentioned a certain illness; when recommending restaurants, the AI notes "this one is too spicy, you previously..." — reasonable-sounding, but the user may feel the AI over-exploited their privacy. Say too much and it feels intrusive; say too little and there's no presence. The demand proactive service places on "sensing the right degree" may exceed what current base models can solve. Zhu's judgment: this may need to exist at the base-model level — not fixable via post-training or PE.
3. The evolution from Harness Engineering to AIWare Engineering, grounded in computing history.
Zhu mapped Agent engineering onto computer history in four stages:
| Year | Computing-history event | Agent engineering stage | |---|---|---| | 1946 | ENIAC | Large model (raw compute) | | 1949 | Assembly language | Prompt Engineering | | 1968 | NATO software crisis conference ("software engineering" coined) | AIWare Engineering | | 1969 | Unix operating system | Agent frameworks / Harness | | 1972 | C language | Context Engineering |
The key insight: 1968 was the real turning point. CPUs, operating systems, and C existed, and people thought the software problem was solved — only to find that software's real difficulty "isn't code, it's people: expectations, project management, team organization." Transposed to AI: the next step beyond Harness Engineering is AIWare Engineering — how to truly integrate humans and AI.
4. "Low power, good enough" is the most contrarian judgment of the talk.
Zhu cited human evolution: language and brain capacity grew 70,000 years ago; afterwards writing, paper, printing, radio, the internet — but brain capacity never grew again. Those advances were essentially optimizing "context engineering": storing, transmitting, and speeding information flow and collaboration. Applied to AI: models could be more precise, but is the ROI worth it? Is the energy and cost demand justified? Perhaps lower-power models with better harness methods can roughly match stronger models — and for society overall, that's more appropriate. Base model + Harness should converge to a reasonable equilibrium, not endlessly maximize a single dimension.
Why It Matters
This is a rare C-end Agent retrospective combining macro vision with hard engineering numbers. Three signals:
Signal one: 1/3 the time + 1/10 the tokens — the cost curve for C-end agents has been pushed to the limit by a Chinese major. This is a key reference for global C-end agent commercialization competition over the next 12–18 months.
Signal two: "Proactive service" will be the main battlefield for C-end agents in H2 2026. A 300M-MAU app pushing this into product form means competitors must answer "how will you proactively push?" head-on.
Signal three: "AIWare Engineering" is becoming a new buzzword. The path from Prompt → Context → Harness → AIWare will be repeatedly cited in upcoming talks, podcasts, and startup camps — a concept mine for the next six months.
Risks and What to Watch
First, Zhu's observation that "many previous scaffolds become unnecessary as base models improve" carries a lesson: many current harness engineering patches will be absorbed by the next generation of base models, so the moat of harness engineering may be shorter than assumed.
Second, "emotional intelligence" — the hardest problem — has no systematic solution from any model or team yet. Qwen's current proactive service is likely still in a "fine-grained probing" stage rather than a full capability breakthrough. Watching the Qwen App's actual behavior over the next 3–6 months will be more informative than the talk itself.
Third, "low-power, good enough" is an engineering philosophy, not a business strategy. If other majors (especially the OpenAI/Anthropic camp) keep "higher intelligence, stronger base models" as their competitive mainline, whether Qwen's low-power philosophy can withstand the base-model gap remains to be seen.
Source: keynote by Zhu Da, head of Alibaba ATH-Qwen MOS Lab, at the CCF YOCSEF Hangzhou technical forum (June 7, 2026).