English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News | November 24, 2025

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for November 24, 2025 covers a packed day in AI: Anthropic released Claude Opus 4.5, priced 3x lower than Opus 4.1 ($5/$25 per million tokens), scoring a SOTA 80.9% on SWE-Bench Verified with new effort control and context compaction APIs. Google launched Gemini 3 Pro (76.2% SWE-Bench Verified, though users report hallucinations and instruction-following issues) and Gemini 3 Image, which topped Artificial Analysis image benchmarks. OpenAI shipped GPT-5.1-Codex-Max at 77.9% SWE-Bench Verified. Zyphra unveiled ZAYA1-base, an AMD-native MoE model (8.3B total, 760M active) beating Llama-3-8B. Other news: Windsurf supports Opus 4.5, Weaviate v1.32 defaults to 8-bit Rotational Quantization, Weights & Biases launched Serverless LoRA inference, Sakana AI's Continuous Thought Machines NeurIPS spotlight, the White House Genesis Mission for AI-for-science, Qwen3-Next llama.cpp support, and community debates on Manus.im's Chat Mode removal.

Easy AI Daily News | November 24, 2025

A daily digest of AI industry updates compiled by the Easy AI teaching project.

Model Updates and Performance

  • Anthropic releases Claude Opus 4.5 — Positioned as the best model for coding and agentic tasks. Priced 3x lower than Opus 4.1 at $5/$25 per million tokens, with new API features including effort control. Scores 80.9% on SWE-Bench Verified, a new SOTA. (official release | Twitter announcement)
  • Google releases Gemini 3 Pro — Achieves 76.2% on SWE-Bench Verified, but user feedback reports hallucination issues and frequent instruction-ignoring. (benchmark results)
  • OpenAI releases GPT-5.1-Codex-Max — Scores 77.9% on SWE-Bench Verified, briefly holding SOTA with improved coding capability.
  • Google releases Gemini 3 Image — Tops the Artificial Analysis image benchmark, supports 14 input images, with improved photorealism and editing. (official tweet)
  • Benchmarks and Reasoning

  • Claude Opus 4.5 sets multiple records: SWE-Bench Verified 80.9%, SWE-bench Pro 52%, BrowseComp-Plus 85.3%, ARC-AGI-1 80%, ARC-AGI-2 37.64%. (system card)
  • Qwen3-VL-32B beats Kimi K1.5 on the MathVision benchmark by 24.8 points, showing stronger visual reasoning.
  • Tools and Ecosystem

  • Claude Opus 4.5 API additions: effort control, context compaction, and advanced tool use; available on Bedrock, Vertex, and other cloud platforms. (effort docs)
  • Windsurf 1.12.35 (stable) / 1.12.152 (preview) adds Claude Opus 4.5 support, offered at Sonnet pricing (2x credits) for a limited time. (changelog)
  • Weaviate v1.32 enables 8-bit Rotational Quantization by default, claiming 98-99% accuracy retention with lower latency and better write performance.
  • Weights & Biases launches Serverless LoRA inference — upload adapters and switch dynamically at inference time with no cold starts.
  • Research and Technical Breakthroughs

  • Zyphra launches ZAYA1-base, an AMD-native MoE model built with AMD and IBM (8.3B total / 760M active parameters) that outperforms Llama-3-8B on math and coding. (technical details)
  • DiRL framework combines SFT with the diffusion-native RL algorithm DiPO to optimize diffusion language models; an 8B model reaches 83% on MATH500.
  • Sakana AI's Continuous Thought Machines (CTM), a NeurIPS spotlight, uses neuron-level dynamics and synchronization for adaptive computation and emergent sequential reasoning, excelling at maze planning.
  • Industry and Policy

  • White House launches the Genesis Mission to accelerate scientific discovery with AI; Anthropic partners with the US Department of Energy. (Anthropic announcement)
  • Community and Product Feedback

  • Gemini 3 users report severe hallucinations and frequent disregard of explicit instructions (e.g., generating a third option when told not to).
  • Manus.im users protest removal of Chat Mode and forced switch to Agent Mode.
  • An Anthropic engineer claimed software engineering will be "finished" in the first half of next year, with AI-generated code as trusted as compiler output, sparking debate. (Reddit discussion)
  • LM Studio users request deprecating an unused system prompt section.
  • Open Source and Local Models

  • ArliAI releases GLM-4.5-Air-Derestricted, using Norm-Preserving Biprojected Abliteration to remove refusal behavior while preserving reasoning. (Hugging Face)
  • Qwen3-Next models now supported in llama.cpp; a user reports 12 tokens/sec on an RTX 5070 Ti. (GitHub PR)
  • Safety and Jailbreaks

  • The BASI Jailbreaking community published a Gemini 3.0 jailbreak guide that bypasses safety filters by uploading Google Docs files, with multilingual prompts (e.g., Croatian). (jailbreak guide)
---

*Source: Easy AI teaching project.*

Tags

#ai-news#claude-opus-4-5#gemini-3#gpt-5-1-codex-max#swe-bench#moe-models#open-source-llms#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169138