Summary
A paper by Peng Gang (arXiv:2603.11113) examines how reliably structured intent representations preserve user goals across AI models, languages, and prompting frameworks. Building on PPS (Prompt Protocol Specification), a 5W3H-based structured intent framework previously validated in Chinese, English, and Japanese, the study extends the work in three directions: cross-model robustness across Claude, GPT-4o, and Gemini 2.5 Pro; controlled comparison with CO-STAR and RISEN frameworks; and a user study (N=50) of AI-assisted intent expansion. Across 3,240 model outputs (3 languages x 6 conditions x 3 models x 3 domains x 20 tasks) judged by an independent evaluator (DeepSeek-V3), structured prompting substantially reduced cross-language score variance, with the strongest structured condition lowering cross-language sigma from 0.470 to about 0.020. The authors also observed a weak-model compensation pattern: Gemini, the weakest baseline model, showed a much larger D-A gain (+1.006) than Claude (+0.217). At current evaluation resolution, 5W3H, CO-STAR, and RISEN achieved similar high goal-alignment scores, suggesting dimensional decomposition itself is a key active ingredient.
Paper Overview
Field: AI
Author: Peng Gang
Published: 2026-03-31
arXiv: 2603.11113
Abstract
How reliably can structured intent representations preserve user goals across different AI models, languages, and prompting frameworks? Prior work showed that PPS (Prompt Protocol Specification), a 5W3H-based structured intent framework, improves goal alignment in Chinese and generalizes to English and Japanese. This paper extends that line of inquiry in three directions:
- Cross-model robustness across Claude, GPT-4o, and Gemini 2.5 Pro
- Controlled comparison with CO-STAR and RISEN prompting frameworks
- User study (N=50) of AI-assisted intent expansion in ecologically valid settings
Key Findings
- Evaluated across 3,240 model outputs (3 languages x 6 conditions x 3 models x 3 domains x 20 tasks), judged by an independent evaluator (DeepSeek-V3).
- Structured prompting substantially reduces cross-language score variance compared to unstructured baselines.
- The strongest structured condition lowered cross-language sigma from 0.470 to about 0.020.
- Weak-model compensation pattern: Gemini, the weakest baseline model, showed a much larger D-A gain (+1.006) than the strongest model Claude (+0.217).
- At current evaluation resolution, 5W3H, CO-STAR, and RISEN achieved similar high goal-alignment scores, indicating that dimensional decomposition itself is an important active ingredient.
---
*Originally posted on zhichai.net; auto-collected 2026-04-02.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169492