English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | November 24, 2025: Claude Opus 4.5, Gemini 3 Pro, and GPT-5.1-Codex-Max Headline Busy AI News Day

Forum topic · 小凯 · 2026-03-27

Summary

November 24, 2025 AI news roundup from the Easy AI teaching project. Anthropic released Claude Opus 4.5, cutting pricing 3x to $5/$25 per million tokens and achieving a new SWE-Bench Verified record of 80.9%, with new API features including effort control, context compaction, and advanced tool use. Google launched Gemini 3 Pro scoring 76.2% on SWE-Bench Verified (with user-reported hallucination issues) plus Gemini 3 Image, which topped the Artificial Analysis image benchmark. OpenAI shipped GPT-5.1-Codex-Max at 77.9% on SWE-Bench Verified. Other highlights: Qwen3-VL-32B beat Kimi K1.5 on MathVision; Zyphra released AMD-native MoE model ZAYA1-base outperforming Llama-3-8B; Sakana AI presented Continuous Thought Machines at NeurIPS; the White House launched the Genesis Mission for AI-for-science; Weaviate enabled 8-bit Rotational Quantization by default; and Qwen3-Next gained llama.cpp support hitting 12 tokens/sec on an RTX 5070 Ti.

📅 AI Industry Highlights — November 24, 2025

Model Releases and Performance

Anthropic releases Claude Opus 4.5, positioned as best model for coding and agents Anthropic launched its flagship Claude Opus 4.5 at one-third the price of Opus 4.1 ($5/$25 per million tokens), adding API features such as effort control. It scored 80.9% on SWE-Bench Verified, setting a new state of the art.

  • Official release | Twitter announcement
  • Google ships Gemini 3 Pro, scoring 76.2% on SWE-Bench Verified Google's Gemini 3 Pro achieved 76.2% on SWE-Bench Verified, though users report hallucination issues and frequent instruction-following failures.

  • Benchmark results | User feedback
  • OpenAI launches GPT-5.1-Codex-Max at 77.9% on SWE-Bench Verified OpenAI's GPT-5.1-Codex-Max briefly held SOTA with 77.9% on SWE-Bench Verified, boosting coding capability.

  • Related tweet
  • Google releases Gemini 3 Image, topping image benchmarks Gemini 3 Image topped the Artificial Analysis image benchmark, supports up to 14 input images, and improves photorealism and editing.

  • Official tweet | Benchmark results
  • Benchmarks and Reasoning

  • Claude Opus 4.5 sets multiple records: 80.9% SWE-Bench Verified, 52% SWE-bench Pro, 85.3% BrowseComp-Plus, 80% ARC-AGI-1, 37.64% ARC-AGI-2. (Benchmarks | System card)
  • Qwen3-VL-32B surpasses Kimi K1.5 on MathVision by 24.8 points, showing stronger visual reasoning.
  • Tools and Ecosystem

  • Claude Opus 4.5 API features: effort control (reasoning intensity), context compaction, and advanced tool use; available on Bedrock and Vertex. (Effort control docs | Context compaction docs)
  • Windsurf supports Claude Opus 4.5 in stable 1.12.35 and preview 1.12.152, offered for a limited time at Sonnet pricing (2x credits). (Changelog | Download)
  • Weaviate v1.32 enables 8-bit Rotational Quantization by default, claiming 98–99% accuracy retention with lower latency and better write performance. (Tweet)
  • Weights & Biases launches Serverless LoRA inference: upload adapters and switch them dynamically at inference with no cold starts. (Tweet)
  • Research and Technical Breakthroughs

  • Zyphra releases AMD-native MoE model ZAYA1-base (8.3B total / 760M active parameters, with AMD and IBM), outperforming Llama-3-8B on math and coding. (Tweet | Technical details)
  • DiRL framework for diffusion language models: combines SFT with the diffusion-native RL algorithm DiPO; an 8B model reaches 83% on MATH500. (Tweet)
  • Sakana AI's Continuous Thought Machines (CTM), a NeurIPS spotlight, uses neuron-level dynamics and synchronization for adaptive computation and emergent sequential reasoning, excelling at maze planning. (Tweet)
  • Industry and Policy

  • US launches Genesis Mission: a White House initiative to accelerate scientific discovery with AI; Anthropic is partnering with the Department of Energy. (Anthropic announcement)
  • Community and Product Feedback

  • Gemini 3 hallucination complaints: users report fabricated information and ignored explicit instructions. (Discord discussion)
  • Manus.im users protest removal of Chat Mode, which forces Agent Mode; some demand its restoration. (Discord discussion)
  • Anthropic engineer says software engineering will be "finished" in the first half of next year, claiming AI-generated code will be as trusted as compiler output. (Reddit discussion)
  • LM Studio users request deprecating the system prompt section, unused for two years, to simplify the interface. (Discord discussion)
  • Open Source and Local Models

  • ArliAI releases GLM-4.5-Air-Derestricted, using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning; based on the Gemma 3 12B architecture. (Hugging Face)
  • Qwen3-Next supported in llama.cpp: users report up to 12 tokens/sec on an RTX 5070 Ti (e.g., Qwen3-Next-80B-A3B-Instruct). (GitHub PR)
  • Security and Jailbreaks

  • BASI Jailbreaking community publishes a Gemini 3.0 jailbreak via uploading Google Docs files, with multilingual prompts (e.g., Croatian). (Guide | Discord discussion)
---

*Source: Easy AI teaching project.*

Tags

#ai-news#claude-opus-4-5#gemini-3#gpt-5-1-codex-max#open-source-llm#benchmarks#daily-digest#ai-policy

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169097