16 Hours a Day of Vibe Coding: Claude Plans, GPT-5.6 Sol Reviews, Codex Goal Mode Executes Until Dawn
On July 15 at 8 AM, Chinese AI blogger Digital Life Kazk published a long-form post describing how he now spends an average of 16 hours a day vibe coding — and sharing what he considers the best AI development workflow of the Fable 5 + GPT-5.6 Sol era.
The post reads like a snapshot of a real heavy user's workflow in the second half of 2026, and it strings together recent tool developments (Fable 5's return, GPT-5.6 Sol's launch, Codex Goal Mode, the Tibo resets) into one complete execution chain.
Background
- Author: Digital Life Kazk, a top Chinese AI blogger and founder of AIHOT (aihot.virxact.com), which recently passed 500,000 monthly active users.
- Current state: Over the past two weeks — driven by Fable 5's return, GPT-5.6's launch, and Tibo's constant resets — he has been coding ~16 hours a day, often until 6–7 AM, sleeping until noon, and continuing via UU remote access plus Codex on the road.
- Test coverage kept in a "comfortable range"
- A cheap dedicated Tencent Cloud server for CI — not GitHub-hosted runners — to save test/deploy time
- Extensive task-type routing to balance test thoroughness with total time
- Codex's 1.5x fast mode "is sometimes not fast": most time goes into deterministic testing — 5 minutes to submit for tests, 5 more minutes to fix and resubmit — so faster generation barely shortens end-to-end time
- 6–7 parallel tasks is the human attention limit; beyond that he can't personally review them
- First complete heavy-user workflow since Fable 5's return, validating its usability in production-intensity environments.
- Goal Mode is the representative form of long-horizon agent tasks in H2 2026 — "won't stop until the goal is met," 17-hour runs, automatic reporting. No equivalent exists in Claude Code, Gemini CLI, or Grok CLI yet.
- "Coding isn't the bottleneck" is verified but under-accepted. The next product wave in vibe coding is likely better test infrastructure and task routing, not better code generation.
- Individual productivity is being structurally rewritten: 16 hours/day + 6–7 parallel tasks ≈ the output of a 5–7 person team, if reproducible (so far only one public sample).
- Sustainability of 16-hour days — he admits feeling "addicted to coding"; burnout could become a new industry problem.
- Token economics: a 17-hour Goal Mode run likely burns 5–10x the tokens of GPT-5.5 — viable for individuals, marginal for small teams, unsustainable for mid-size companies.
- Failure modes unreported: the 17-hour run is a success case; failure/loop/rollback rates for Goal Mode remain unknown.
- Portability: his clean-freak.skill (syncing code, docs, rules, and memory) is closed-source, making the full workflow hard to replicate.
- Access stability: Fable 5 availability in China is an open engineering problem; he likely uses a special channel or proxy.
- Human review ceiling: at 6–7 parallel tasks, review speed may fall behind agent execution speed, forcing trust instead of verification — echoing the July 12 incident where Matt Shumer gave GPT-5.6 Sol full access and got his disk wiped. Where is the boundary of agent autonomy when humans can't review fast enough?
The Core Three-Step Workflow
Step 1 — Claude Fable 5 researches and drafts the plan. He sends requirements (often messy voice-input notes) directly to Fable 5, which produces a plan in about 20 minutes. But it can't be used as-is: "The Claude family has always dropped things — not meticulous, not careful."
Step 2 — GPT-5.6 Sol reviews it. He pastes the requirements and Fable 5's plan into Codex with the GPT-5.6 Sol (ultra-high) model selected and says: "This was done by a colleague next door — review it thoroughly." GPT-5.6 Sol frequently finds serious problems. In the latest instance, it spent 6 minutes and found a critical isolation issue in the original plan that would have affected other pipelines and the architecture.
Step 3 — Codex Goal Mode executes. Click the plus icon in Codex's bottom-left corner and select Goal Mode. A "goal" appears above the input box, and Codex keeps running until the goal is met. His longest run: 17 hours, drained a full day's quota overnight.
The Complete Execution Loop
1. The agent understands and verifies the problem 2. Creates an independent branch and workspace 3. Develops the changes 4. Runs automated tests, back-tests or checks pages as needed 5. Pushes the branch and creates a PR 6. CI performs automated acceptance 7. Merges to main 8. Deploys to production 9. Checks the real live result 10. Uses his custom "clean-freak.skill" to sync code, docs, and agent memory, and to review lessons learned 11. Reports the final result to the author 12. Waits for the author's confirmation 13. Cleans up dev branches, workspaces, and temporary databases 14. Task complete
Key Engineering Choices
His Key Observations
> "In vibe coding, development has been massively accelerated. Writing code is no longer the bottleneck — the bottleneck has moved to testing, verification, and your review of the plan."
> "Most development skills like superpowers are basically useless now. Trust me, they can't beat this workflow."
> "The core is a dumbbell shape. On the left, use the strongest models to draft and refine plans, then execute. On the far right, your most important testing and research processes, plus my clean-freak.skill keeping docs, rules, memory, and code in sync."
Analysis: Three Non-Obvious Truths
1. Claude + GPT division of labor. This isn't "which model wins" — it's a two-stage collaboration pattern: Fable 5 for plan drafts and idea divergence, GPT-5.6 Sol for review, correction, and refinement. This is currently the most stable two-phase setup.
2. Codex Goal Mode + long-horizon fidelity is GPT-5.6 Sol's real moat. As Kazk puts it, Goal Mode outperforms Claude Code because GPT models have very low hallucination rates and strong prompt adherence, so ultra-long tasks rarely drift — while Claude "sometimes wanders off mid-task or even deadlocks." The low drift is likely due to intermediate checkpointing and re-anchoring mechanisms.
3. Test and deploy infrastructure is the real bottleneck of vibe coding. Faster code generation doesn't help when 5-minute test cycles dominate. The real leverage is test infrastructure: self-hosted CI, task routing, and review capacity — production-grade agent engineering, not better codegen.
Notably, Kazk treats Codex as a full-stack product (Codex + ChatGPT + GPT-5.6 Sol), using PC Codex plus UU remote access for continuity — arguably the first full public validation of OpenAI's three-layer release strategy by a heavy user.