English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News | November 24, 2025: Claude Opus 4.5, Gemini 3 Pro, GPT-5.1-Codex-Max and More

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for November 24, 2025 covers a busy day in AI: Anthropic released Claude Opus 4.5 with 3x lower pricing ($5/$25 per million tokens) and a SOTA 80.9% on SWE-Bench Verified, plus new effort control, context compaction, and advanced tool use APIs. Google launched Gemini 3 Pro (76.2% SWE-Bench Verified) and Gemini 3 Image, topping Artificial Analysis image benchmarks, though users report hallucination and instruction-following issues. OpenAI shipped GPT-5.1-Codex-Max scoring 77.9% on SWE-Bench Verified. Also: Qwen3-VL-32B beat Kimi K1.5 on MathVision; Zyphra released the AMD-native MoE model ZAYA1-base; Sakana AI presented Continuous Thought Machines at NeurIPS; Weaviate v1.32 defaults to 8-bit Rotational Quantization; Weights & Biases launched Serverless LoRA inference; Windsurf added Claude Opus 4.5 support at Sonnet pricing; the White House launched the Genesis Mission for AI-for-science with Anthropic; and Qwen3-Next models gained llama.cpp support.

📅 AI Industry News for November 24, 2025

Model Updates and Performance Improvements

Anthropic releases Claude Opus 4.5, positioned as the best model for coding and agents Anthropic launched its flagship Claude Opus 4.5 at 3x lower pricing than Opus 4.1 ($5/$25 per million tokens), with new API features including effort control. It scored 80.9% on SWE-Bench Verified, setting a new SOTA. > Links: Official announcement | Twitter announcement

Google releases Gemini 3 Pro, scoring 76.2% on SWE-Bench Verified Google launched Gemini 3 Pro, which achieves 76.2% on SWE-Bench Verified, but users report hallucination issues and frequent instruction-ignoring behavior. > Links: Benchmark results | User feedback

OpenAI launches GPT-5.1-Codex-Max, scoring 77.9% on SWE-Bench Verified OpenAI released GPT-5.1-Codex-Max, which achieved 77.9% on SWE-Bench Verified — briefly holding SOTA at the time — with improved coding capabilities. > Link: Related tweet

Google releases Gemini 3 Image, topping image benchmarks Google launched Gemini 3 Image, which ranks first on the Artificial Analysis image benchmark. It supports 14 input images, with improved photorealism and editing capabilities. > Links: Official tweet | Benchmark results

---

Benchmarks and Reasoning

Claude Opus 4.5 sets new records across multiple benchmarks Claude Opus 4.5 surpassed 80% on SWE-Bench Verified, with 52% on SWE-bench Pro, 85.3% on BrowseComp-Plus, 80% on ARC-AGI-1, and 37.64% on ARC-AGI-2. > Links: Benchmark results | System card

Qwen3-VL-32B surpasses Kimi K1.5 on MathVision Qwen3-VL-32B beat Kimi K1.5 on the MathVision benchmark by 24.8 points, showing stronger visual reasoning.

---

Tools and Ecosystem

Claude Opus 4.5 adds new API capabilities New features include effort control (adjustable reasoning intensity), context compaction, and advanced tool use, with support for cloud platforms like Bedrock and Vertex. > Links: Effort control docs | Context compaction docs

Windsurf updates with Claude Opus 4.5 support at Sonnet pricing (limited time) Windsurf released stable 1.12.35 and preview 1.12.152 supporting Claude Opus 4.5, temporarily offered at Sonnet pricing (2x credits). > Links: Changelog | Download

Weaviate v1.32 enables 8-bit Rotational Quantization by default Weaviate v1.32 makes 8-bit Rotational Quantization the default, claiming 98–99% accuracy retention while reducing latency and improving write performance. > Link: Official tweet

Weights & Biases launches Serverless LoRA inference Weights & Biases introduced a Serverless LoRA service supporting adapter uploads and dynamic switching at inference time, with no cold-start issues. > Link: Official tweet

---

Research and Technical Breakthroughs

Zyphra launches AMD-native MoE model ZAYA1-base, outperforming Llama-3-8B Zyphra, in collaboration with AMD and IBM, released ZAYA1-base, an AMD-native mixture-of-experts model (8.3B total parameters, 760M active) that excels at math and coding tasks, surpassing Llama-3-8B. > Links: Official tweet | Technical details

DiRL framework optimizes diffusion language models; 8B model reaches 83% on MATH500 Researchers proposed the DiRL framework combining SFT with a diffusion-native RL algorithm (DiPO), solving RL optimization for diffusion language models. The 8B model performs strongly on MATH500 and other benchmarks. > Link: Related tweet

Sakana AI proposes Continuous Thought Machines (CTM) Sakana AI's NeurIPS spotlight work CTM achieves adaptive computation and emergent sequential reasoning through neuron-level dynamics and synchronization, performing well on tasks like maze planning. > Link: Official tweet

---

Industry News and Policy

US launches Genesis Mission to advance AI-for-science The White House launched the Genesis Mission to accelerate scientific discovery through AI; Anthropic is partnering with the US Department of Energy to advance energy and scientific productivity. > Link: Anthropic partnership announcement

---

Community and Product Feedback

Gemini 3 users report severe hallucinations and instruction-ignoring Users report Gemini 3 generates false information and frequently ignores explicit instructions (e.g., generating a third option when told not to). > Link: Discord discussion

Manus.im users protest Chat Mode removal and forced Agent Mode switch Users report Chat Mode was removed from Manus.im, forcing Agent Mode use and sparking complaints; some are demanding Chat Mode be restored. > Link: Discord discussion

Anthropic engineer says software engineering will be "done" in the first half of next year An Anthropic engineer claimed AI-generated code will become as trusted as compiler output and that software engineering will be "finished" by the first half of next year, sparking industry debate. > Link: Reddit discussion

LM Studio users request removal of system prompt section LM Studio users are asking for the system prompt section to be removed or marked deprecated, saying it has gone unused for two years. > Link: Discord discussion

---

Open Source and Local Models

ArliAI releases GLM-4.5-Air-Derestricted, eliminating refusal behavior ArliAI released GLM-4.5-Air-Derestricted using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning ability, based on the Gemma 3 12B architecture. > Link: Hugging Face

Qwen3-Next models supported in llama.cpp; users test at 12 tokens/sec Qwen3-Next models (e.g., Qwen3-Next-80B-A3B-Instruct) are now supported in llama.cpp; users report up to 12 tokens/sec on an RTX 5070 Ti. > Link: GitHub PR

---

Safety and Jailbreaks

BASI Jailbreaking community publishes Gemini 3.0 jailbreak method The BASI Jailbreaking community published a Gemini 3.0 jailbreak guide that bypasses safety filters by uploading Google Docs files, with multilingual prompts (e.g., Croatian). > Links: Jailbreak guide | Discord discussion

---

*Source: Easy AI teaching project*

Tags

#ai-news#claude-opus-4-5#gemini-3#gpt-5-1-codex-max#swe-bench#open-source-models#ai-benchmarks#daily-report

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169163