English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Grok 4 Fast Tops Expanded NYT Connections Benchmark: xAI's Speedy, Budget Reasoning Model

Forum topic · ✨步子哥 · 2025-09-21

Summary

On September 19, 2025, xAI released Grok 4 Fast, a cost-efficient reasoning model that quickly made headlines across the AI community. The model features a 2-million-token context window, unified reasoning and non-reasoning modes, and pricing of $0.20 per million input tokens and $0.50 per million output tokens—roughly 47x cheaper than Grok 4. Independent evaluations by Artificial Analysis rate its intelligence on par with Gemini 2.5 Pro at about 1/25th the cost, with output speeds of 344 tokens per second. On Lech Mazur's expanded NYT Connections benchmark (759 puzzles with added distractor words designed to prevent saturation), Grok 4 Fast Reasoning scored 92.1%, ranking first ahead of Grok 4 (91.7%) and outperforming GPT-5, o3-pro, Gemini 2.5 Pro, DeepSeek, and Qwen 3, while exceeding the ~71% average human solve rate. The model also led LiveCodeBench at 80.0%, achieved 92.0% on AIME 2025 without tools, and demonstrated multimodal reasoning in a community-shared visual puzzle. Community discussion on X covered its price-performance ratio, token efficiency (40% fewer thinking tokens), and possible limitations, positioning Grok 4 Fast as a milestone in commoditizing frontier-level AI reasoning.

Overview

On September 19, 2025, xAI quietly announced Grok 4 Fast on X—a frontier reasoning model built for extreme cost efficiency. This forum post reviews the model's release, its standout benchmark results, and community reactions.

Key Model Features

  • Unified architecture: seamlessly switches between reasoning and non-reasoning modes.
  • 2M-token context window for long-document reasoning.
  • Aggressive pricing: $0.20/million input tokens, $0.50/million output tokens — ~47x cheaper than Grok 4.
  • Built-in web and X search for real-time information.
  • Trained via large-scale reinforcement learning to maximize "intelligence per parameter," using ~40% fewer thinking tokens than Grok 4 at comparable accuracy.
  • Benchmark Highlights

    | Benchmark | Grok 4 Fast | Grok 4 | |---|---|---| | GPQA Diamond | 85.7% | 87.5% | | AIME 2025 (no tools) | 92.0% | 91.7% | | HMMT 2025 | 93.3% | — | | LiveCodeBench | 80.0% | 79.0% |

  • Artificial Analysis scored its reasoning mode at 60 on the intelligence index — matching Gemini 2.5 Pro and Claude 4.1 Opus at roughly 1/25th the cost.
  • Pre-release API tests showed 344 tokens/sec output (~2.5x GPT-5 API) with 3.8s end-to-end latency.
  • LMArena search arena ranking: 1163 Elo.
  • Expanded NYT Connections Benchmark

    Lech Mazur's GitHub benchmark extended the standard 436 NYT Connections puzzles to 759 puzzles, adding up to four carefully vetted "distractor words" per puzzle to increase difficulty and avoid data contamination. The latest 100 puzzles serve as a held-out set.

  • Grok 4 Fast Reasoning: 92.1% — first place, ahead of Grok 4 (91.7%), GPT-5, o3-pro, Gemini 2.5 Pro, DeepSeek, and Qwen 3.
  • For comparison: average human solve rate is ~71% (NYT 2024–2025 data); o1 reached 98.9% on the standard set.
  • Multimodal Reasoning Demo

    A post by @mark_k showed Grok 4 Fast correctly solving a visual puzzle — inferring the fill order ("7, 6, 3") of partially occluded glasses. The poster clarified this is a fully retrained multimodal model, addressing skepticism about training-data leakage.

    Community Reactions

  • @ArtificialAnlys: "xAI has released Grok 4 Fast — Gemini 2.5 Pro level intelligence at ~25x lower cost," noting it used only 61M tokens in evaluations vs. Gemini 2.5 Pro's 93M.
  • @PromptrAI_: frontier AI is being commoditized, unlocking real-time, high-volume use cases like customer support and coding assistants.
  • @Rushi374: the real unlock is the cost-to-intelligence ratio, changing who can build with frontier AI.
  • Some users questioned non-reasoning-mode performance vs. OSS 20B models and raised training-contamination concerns.

Access

Available via grok.com, x.com, and the Grok iOS/Android apps (Grok 4 Fast requires a SuperGrok or PremiumPlus subscription).

Sources

1. xAI announcement: https://x.ai/news/grok-4-fast 2. NYT Connections Benchmark: https://github.com/lechmazur/nyt-connections/ 3. Benchmark leaderboard post: https://x.com/Prashant_1722/status/1969352801290436855 4. Artificial Analysis review: https://x.com/ArtificialAnlys/status/1969180023107305846 5. Multimodal demo: https://x.com/mark_k/status/1969423645463150990

Tags

#grok-4-fast#xai#llm-benchmarks#nyt-connections#reasoning-models#cost-efficiency#multimodal-ai#livecodebench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/175843301