English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | December 9, 2025: GLM-4.6V, Hugging Face Training Skills, Stargate DRAM Shortage, and More

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for December 9, 2025 covers major AI industry updates. Zhipu released GLM-4.6V (106B MoE) and GLM-4.6V-Flash (9B dense) multimodal models with 128k context and native function calling, with Flash API free and weights on Hugging Face. JinaAI launched Jina-VLM-2B, achieving SOTA on multilingual VQA benchmarks. Hugging Face introduced Claude Code skills for natural-language LLM training automation (~$0.30 per small run). Google's NeurIPS Miras framework outperforms Transformer/Mamba2 on long-context tasks. AxiomProver solved 9/12 Putnam problems using Lean formal verification. OpenAI's Stargate project will consume up to 40% of global DRAM output, raising DDR5 prices. Also featured: LangChain Deep Agents evaluation, AMD 7900XTX as budget local LLM GPU, RadixArk spinoff from SGLang, Meta's acquisition of Limitless, and ARC Prize 2025 results where NVARC scored 25.03% with the $600k grand prize unclaimed.

Easy AI Daily | December 9, 2025

*Translated from the zhichai.net Easy AI Daily digest.*

Model Updates & Multimodal Progress

  • Zhipu releases GLM-4.6V and GLM-4.6V-Flash: GLM-4.6V is a 106B MoE model for cloud/high-performance clusters, while GLM-4.6V-Flash is a 9B dense model for local, low-latency use. Both support 128k context and native multimodal function calling. The Flash API is free, and weights are open on Hugging Face. (GLM-4.6V-Flash | GLM-4.6V)
  • JinaAI releases Jina-VLM-2B: A 2B-parameter multilingual VLM focused on charts, documents, and scene text. It averages 72.3 across 8 VQA benchmarks, with SOTA on MMMB (78.8) and multilingual MMBench (74.3). (Announcement)
  • Qwen3-4B shines locally: Reaches 70 tokens/sec on an RTX 2060 with strong coding ability, though it tends to "overthink"; users recommend pairing it with GLM-4.6V-Flash.
  • DeepSeek V3.2 improves reasoning: Supports interleaved reasoning with strong RoO code performance. Pricing: $0.28/M input tokens, $0.45/M output tokens; reportedly more stable than Kimi.
  • Training & Optimization Tools

  • Hugging Face launches Claude Code skills for LLM training automation: Specify training tasks in natural language (e.g., fine-tuning Qwen3-0.6B); the system handles data validation, GPU selection, and HF Jobs submission. Small runs cost around $0.30. (Blog)
  • Unsloth AI progress: Celebrated 10k Reddit members, fixed slow HF downloads, and released Mistral Large 3 GGUF models for local inference. (Reddit | GGUF)
  • DSPy TOON Adapter: Reduces token consumption, though it handles nested schemas less well than BAMLAdapter; MMLU-Pro improves after GEPA optimization. (Code)
  • LangChain Deep Agents evaluation: A framework for evaluating long-running agents (planning, filesystem, sub-agents), averaging 42.65% on Terminal Bench 2.0, with context-compression triggers added. (Announcement)
  • Research & Evaluation

  • Google's Miras post-Transformer framework: A NeurIPS paper frames Transformers/RNNs as associative memory systems, outperforming Transformer/Mamba2/DeltaNet on LM, reasoning, and long-context tasks, with a 20% improvement in long-text retrieval. (Overview)
  • AxiomProver solves 9/12 Putnam problems: The Lean-based system solved 9 problems within hours of the exam, emphasizing verifiability and formal pipelines, exceeding last year's top performance. (Announcement)
  • NeurIPS mechanistic interpretability workshop: Chris Olah reflected on interpretability, arguing the field needs scalable tools rather than neuron-level analysis of single models.
  • MEMTRACK benchmark: Evaluates agent long-term memory via Slack/Linear/git scenarios; GPT-5 scored 60%, showing room for improvement in real tool environments. (Announcement)
  • Agents & Workflows

  • Dexter 2.0: An open-source financial research agent with planning and self-verification, built on LangChain, suited for long-horizon financial analysis. (Demo)
  • AI21 Maestro: Agent orchestration with multi-step planning, built-in verification, proprietary RAG, and execution graphs. (Announcement)
  • Cursor Agent issues: Community reports of failures creating files and infinite loops; approval-button workarounds exist but a permanent fix is needed.
  • Infrastructure & Hardware

  • OpenAI Stargate triggers DRAM shortage: Stargate will consume up to 40% of global DRAM output (~900k wafers/month) via deals with Samsung and SK Hynix, driving up DDR5 prices and affecting even gamers' memory costs. (Tom's Hardware)
  • AMD 7900XTX as budget AI GPU: Great price/performance with solid llama.cpp support for local LLMs; ray tracing lags but AI performance is strong. (Discussion)
  • Blackwell WGMMA incompatibility: Blackwell GPUs lack WGMMA (sm90a-only), requiring kernel rewrites; CUDA 13.1 introduces CUDA Tile to simplify programming. (Blog)
  • RadixArk spins off from SGLang: Focused on AI scheduling, compilers, serving, and training pipelines to make frontier infrastructure more open. (Announcement)
  • Community & Ecosystem

  • Local LLM cluster build: One user assembled 8x RTX 3090 (192GB VRAM) + 64-core EPYC Milan + 250GB RAM (~$8k), running GLM-4.5 Air Q6_K at 49 tokens/sec via llama.cpp. (Details)
  • Vector database selection guide: HNSW (<10M vectors), Turbopuffer (large databases), pgvector (small/local), Chroma (lightweight); the guide criticizes some commercial DBs for inefficiency. (Blog)
  • Core community contributors: Unsloth (fast fine-tuning), mradermacher (automated quantization), Bartowski (curated quants), and TheBloke (base models) recognized by users. (Discussion)
  • picomon AMD GPU monitor: Open-source AMD GPU monitoring from a HF user, more reliable than nvtop at the cost of some precision. (Code)
  • Industry & Market

  • Meta acquires Limitless: Formerly Rewind; Pendant users get 1 year of support, while non-Pendant features will sunset. (Announcement)
  • ARC Prize 2025 results: NVARC took Top Score at 25.03%; the TRM paper won $50k; the $600k Grand Prize went unclaimed. Winning methods will be open-sourced. (Results)
  • Manus.im billing issues: Users report lost credits and auto-renewing subscriptions to 2026, with slow support responses, suspected system bug.
  • Sora 2 regional restrictions: Available in only 7 countries; VPN use violates the ToS and may lead to account bans. (Supported countries)
  • Fun & Memes

  • ChatGPT "napping on the job" screenshots go viral. (Meme)
  • The "AI fixes a bug" loop meme captures developer frustration with AI claiming fixes that don't work. (Video)
  • Ultra-detailed David Duchovny AI image with a 20-item checklist verification. (Tweet)
  • Nano Banana Pro fashion editorial workflow using contact-sheet prompts with camera and styling constraints. (Details)
---

*Source: Easy AI education project*

Tags

#ai-news#glm-4.6v#hugging-face#openai-stargate#langchain#local-llm#multimodal-models#deepseek

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169091