📅 AI Industry Highlights — November 24, 2025
Model Releases and Performance
Anthropic releases Claude Opus 4.5, positioned as best model for coding and agents Anthropic launched its flagship Claude Opus 4.5 at one-third the price of Opus 4.1 ($5/$25 per million tokens), adding API features such as effort control. It scored 80.9% on SWE-Bench Verified, setting a new state of the art.
- Official release | Twitter announcement
- Benchmark results | User feedback
- Related tweet
- Official tweet | Benchmark results
- Claude Opus 4.5 sets multiple records: 80.9% SWE-Bench Verified, 52% SWE-bench Pro, 85.3% BrowseComp-Plus, 80% ARC-AGI-1, 37.64% ARC-AGI-2. (Benchmarks | System card)
- Qwen3-VL-32B surpasses Kimi K1.5 on MathVision by 24.8 points, showing stronger visual reasoning.
- Claude Opus 4.5 API features: effort control (reasoning intensity), context compaction, and advanced tool use; available on Bedrock and Vertex. (Effort control docs | Context compaction docs)
- Windsurf supports Claude Opus 4.5 in stable 1.12.35 and preview 1.12.152, offered for a limited time at Sonnet pricing (2x credits). (Changelog | Download)
- Weaviate v1.32 enables 8-bit Rotational Quantization by default, claiming 98–99% accuracy retention with lower latency and better write performance. (Tweet)
- Weights & Biases launches Serverless LoRA inference: upload adapters and switch them dynamically at inference with no cold starts. (Tweet)
- Zyphra releases AMD-native MoE model ZAYA1-base (8.3B total / 760M active parameters, with AMD and IBM), outperforming Llama-3-8B on math and coding. (Tweet | Technical details)
- DiRL framework for diffusion language models: combines SFT with the diffusion-native RL algorithm DiPO; an 8B model reaches 83% on MATH500. (Tweet)
- Sakana AI's Continuous Thought Machines (CTM), a NeurIPS spotlight, uses neuron-level dynamics and synchronization for adaptive computation and emergent sequential reasoning, excelling at maze planning. (Tweet)
- US launches Genesis Mission: a White House initiative to accelerate scientific discovery with AI; Anthropic is partnering with the Department of Energy. (Anthropic announcement)
- Gemini 3 hallucination complaints: users report fabricated information and ignored explicit instructions. (Discord discussion)
- Manus.im users protest removal of Chat Mode, which forces Agent Mode; some demand its restoration. (Discord discussion)
- Anthropic engineer says software engineering will be "finished" in the first half of next year, claiming AI-generated code will be as trusted as compiler output. (Reddit discussion)
- LM Studio users request deprecating the system prompt section, unused for two years, to simplify the interface. (Discord discussion)
- ArliAI releases GLM-4.5-Air-Derestricted, using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning; based on the Gemma 3 12B architecture. (Hugging Face)
- Qwen3-Next supported in llama.cpp: users report up to 12 tokens/sec on an RTX 5070 Ti (e.g., Qwen3-Next-80B-A3B-Instruct). (GitHub PR)
- BASI Jailbreaking community publishes a Gemini 3.0 jailbreak via uploading Google Docs files, with multilingual prompts (e.g., Croatian). (Guide | Discord discussion)
Google ships Gemini 3 Pro, scoring 76.2% on SWE-Bench Verified Google's Gemini 3 Pro achieved 76.2% on SWE-Bench Verified, though users report hallucination issues and frequent instruction-following failures.
OpenAI launches GPT-5.1-Codex-Max at 77.9% on SWE-Bench Verified OpenAI's GPT-5.1-Codex-Max briefly held SOTA with 77.9% on SWE-Bench Verified, boosting coding capability.
Google releases Gemini 3 Image, topping image benchmarks Gemini 3 Image topped the Artificial Analysis image benchmark, supports up to 14 input images, and improves photorealism and editing.
Benchmarks and Reasoning
Tools and Ecosystem
Research and Technical Breakthroughs
Industry and Policy
Community and Product Feedback
Open Source and Local Models
Security and Jailbreaks
*Source: Easy AI teaching project.*