📅 AI Industry News for November 24, 2025
Model Updates and Performance Improvements
Anthropic releases Claude Opus 4.5, positioned as the best model for coding and agents Anthropic launched its flagship Claude Opus 4.5 at 3x lower pricing than Opus 4.1 ($5/$25 per million tokens), with new API features including effort control. It scored 80.9% on SWE-Bench Verified, setting a new SOTA. > Links: Official announcement | Twitter announcement
Google releases Gemini 3 Pro, scoring 76.2% on SWE-Bench Verified Google launched Gemini 3 Pro, which achieves 76.2% on SWE-Bench Verified, but users report hallucination issues and frequent instruction-ignoring behavior. > Links: Benchmark results | User feedback
OpenAI launches GPT-5.1-Codex-Max, scoring 77.9% on SWE-Bench Verified OpenAI released GPT-5.1-Codex-Max, which achieved 77.9% on SWE-Bench Verified — briefly holding SOTA at the time — with improved coding capabilities. > Link: Related tweet
Google releases Gemini 3 Image, topping image benchmarks Google launched Gemini 3 Image, which ranks first on the Artificial Analysis image benchmark. It supports 14 input images, with improved photorealism and editing capabilities. > Links: Official tweet | Benchmark results
---
Benchmarks and Reasoning
Claude Opus 4.5 sets new records across multiple benchmarks Claude Opus 4.5 surpassed 80% on SWE-Bench Verified, with 52% on SWE-bench Pro, 85.3% on BrowseComp-Plus, 80% on ARC-AGI-1, and 37.64% on ARC-AGI-2. > Links: Benchmark results | System card
Qwen3-VL-32B surpasses Kimi K1.5 on MathVision Qwen3-VL-32B beat Kimi K1.5 on the MathVision benchmark by 24.8 points, showing stronger visual reasoning.
---
Tools and Ecosystem
Claude Opus 4.5 adds new API capabilities New features include effort control (adjustable reasoning intensity), context compaction, and advanced tool use, with support for cloud platforms like Bedrock and Vertex. > Links: Effort control docs | Context compaction docs
Windsurf updates with Claude Opus 4.5 support at Sonnet pricing (limited time) Windsurf released stable 1.12.35 and preview 1.12.152 supporting Claude Opus 4.5, temporarily offered at Sonnet pricing (2x credits). > Links: Changelog | Download
Weaviate v1.32 enables 8-bit Rotational Quantization by default Weaviate v1.32 makes 8-bit Rotational Quantization the default, claiming 98–99% accuracy retention while reducing latency and improving write performance. > Link: Official tweet
Weights & Biases launches Serverless LoRA inference Weights & Biases introduced a Serverless LoRA service supporting adapter uploads and dynamic switching at inference time, with no cold-start issues. > Link: Official tweet
---
Research and Technical Breakthroughs
Zyphra launches AMD-native MoE model ZAYA1-base, outperforming Llama-3-8B Zyphra, in collaboration with AMD and IBM, released ZAYA1-base, an AMD-native mixture-of-experts model (8.3B total parameters, 760M active) that excels at math and coding tasks, surpassing Llama-3-8B. > Links: Official tweet | Technical details
DiRL framework optimizes diffusion language models; 8B model reaches 83% on MATH500 Researchers proposed the DiRL framework combining SFT with a diffusion-native RL algorithm (DiPO), solving RL optimization for diffusion language models. The 8B model performs strongly on MATH500 and other benchmarks. > Link: Related tweet
Sakana AI proposes Continuous Thought Machines (CTM) Sakana AI's NeurIPS spotlight work CTM achieves adaptive computation and emergent sequential reasoning through neuron-level dynamics and synchronization, performing well on tasks like maze planning. > Link: Official tweet
---
Industry News and Policy
US launches Genesis Mission to advance AI-for-science The White House launched the Genesis Mission to accelerate scientific discovery through AI; Anthropic is partnering with the US Department of Energy to advance energy and scientific productivity. > Link: Anthropic partnership announcement
---
Community and Product Feedback
Gemini 3 users report severe hallucinations and instruction-ignoring Users report Gemini 3 generates false information and frequently ignores explicit instructions (e.g., generating a third option when told not to). > Link: Discord discussion
Manus.im users protest Chat Mode removal and forced Agent Mode switch Users report Chat Mode was removed from Manus.im, forcing Agent Mode use and sparking complaints; some are demanding Chat Mode be restored. > Link: Discord discussion
Anthropic engineer says software engineering will be "done" in the first half of next year An Anthropic engineer claimed AI-generated code will become as trusted as compiler output and that software engineering will be "finished" by the first half of next year, sparking industry debate. > Link: Reddit discussion
LM Studio users request removal of system prompt section LM Studio users are asking for the system prompt section to be removed or marked deprecated, saying it has gone unused for two years. > Link: Discord discussion
---
Open Source and Local Models
ArliAI releases GLM-4.5-Air-Derestricted, eliminating refusal behavior ArliAI released GLM-4.5-Air-Derestricted using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning ability, based on the Gemma 3 12B architecture. > Link: Hugging Face
Qwen3-Next models supported in llama.cpp; users test at 12 tokens/sec Qwen3-Next models (e.g., Qwen3-Next-80B-A3B-Instruct) are now supported in llama.cpp; users report up to 12 tokens/sec on an RTX 5070 Ti. > Link: GitHub PR
---
Safety and Jailbreaks
BASI Jailbreaking community publishes Gemini 3.0 jailbreak method The BASI Jailbreaking community published a Gemini 3.0 jailbreak guide that bypasses safety filters by uploading Google Docs files, with multilingual prompts (e.g., Croatian). > Links: Jailbreak guide | Discord discussion
---
*Source: Easy AI teaching project*