Easy AI Daily News | November 24, 2025
A daily digest of AI industry updates compiled by the Easy AI teaching project.
Model Updates and Performance
- Anthropic releases Claude Opus 4.5 — Positioned as the best model for coding and agentic tasks. Priced 3x lower than Opus 4.1 at $5/$25 per million tokens, with new API features including effort control. Scores 80.9% on SWE-Bench Verified, a new SOTA. (official release | Twitter announcement)
- Google releases Gemini 3 Pro — Achieves 76.2% on SWE-Bench Verified, but user feedback reports hallucination issues and frequent instruction-ignoring. (benchmark results)
- OpenAI releases GPT-5.1-Codex-Max — Scores 77.9% on SWE-Bench Verified, briefly holding SOTA with improved coding capability.
- Google releases Gemini 3 Image — Tops the Artificial Analysis image benchmark, supports 14 input images, with improved photorealism and editing. (official tweet)
- Claude Opus 4.5 sets multiple records: SWE-Bench Verified 80.9%, SWE-bench Pro 52%, BrowseComp-Plus 85.3%, ARC-AGI-1 80%, ARC-AGI-2 37.64%. (system card)
- Qwen3-VL-32B beats Kimi K1.5 on the MathVision benchmark by 24.8 points, showing stronger visual reasoning.
- Claude Opus 4.5 API additions: effort control, context compaction, and advanced tool use; available on Bedrock, Vertex, and other cloud platforms. (effort docs)
- Windsurf 1.12.35 (stable) / 1.12.152 (preview) adds Claude Opus 4.5 support, offered at Sonnet pricing (2x credits) for a limited time. (changelog)
- Weaviate v1.32 enables 8-bit Rotational Quantization by default, claiming 98-99% accuracy retention with lower latency and better write performance.
- Weights & Biases launches Serverless LoRA inference — upload adapters and switch dynamically at inference time with no cold starts.
- Zyphra launches ZAYA1-base, an AMD-native MoE model built with AMD and IBM (8.3B total / 760M active parameters) that outperforms Llama-3-8B on math and coding. (technical details)
- DiRL framework combines SFT with the diffusion-native RL algorithm DiPO to optimize diffusion language models; an 8B model reaches 83% on MATH500.
- Sakana AI's Continuous Thought Machines (CTM), a NeurIPS spotlight, uses neuron-level dynamics and synchronization for adaptive computation and emergent sequential reasoning, excelling at maze planning.
- White House launches the Genesis Mission to accelerate scientific discovery with AI; Anthropic partners with the US Department of Energy. (Anthropic announcement)
- Gemini 3 users report severe hallucinations and frequent disregard of explicit instructions (e.g., generating a third option when told not to).
- Manus.im users protest removal of Chat Mode and forced switch to Agent Mode.
- An Anthropic engineer claimed software engineering will be "finished" in the first half of next year, with AI-generated code as trusted as compiler output, sparking debate. (Reddit discussion)
- LM Studio users request deprecating an unused system prompt section.
- ArliAI releases GLM-4.5-Air-Derestricted, using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning. (Hugging Face)
- Qwen3-Next models now supported in llama.cpp; a user reports 12 tokens/sec on an RTX 5070 Ti. (GitHub PR)
- The BASI Jailbreaking community published a Gemini 3.0 jailbreak guide that bypasses safety filters by uploading Google Docs files, with multilingual prompts (e.g., Croatian). (jailbreak guide)
Benchmarks and Reasoning
Tools and Ecosystem
Research and Technical Breakthroughs
Industry and Policy
Community and Product Feedback
Open Source and Local Models
Safety and Jailbreaks
*Source: Easy AI teaching project.*