English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News Digest | December 18, 2025

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for December 18, 2025 covers major AI industry updates. Google launched Gemini 3 Flash with Pro-level reasoning at a quarter of the cost, featuring 1M token context and pricing of $0.50/M input tokens; benchmarks show it beating Gemini 3 Pro on ARC-AGI-2 and SWE-bench Verified. xAI released a Grok Voice Agent API with real-time speech-to-speech across 100+ languages at $0.05/min. Microsoft open-sourced TRELLIS 2-4B, an image-to-3D model, while Apple introduced SHARP for photorealistic 3D Gaussians from single images. QwenLong-L1.5 achieved SOTA long-context reasoning with 4M token support. Tencent unveiled Hunyuan HY World 1.5 for real-time interactive 3D world generation. Additional items include Claude Opus API cost complaints, OpenRouter timeout incidents, Unsloth CLI tooling, Qdrant's Snappy PDF search pipeline, RTX PRO 5000 Blackwell spec leaks, and budget GPU rental pricing from NeoCloudX.

Easy AI Daily | December 18, 2025 AI Industry News

Model Releases and Updates

Google Releases Gemini 3 Flash: Pro-Level Reasoning at 1/4 the Cost

Gemini 3 Flash is positioned as "Pro-level reasoning at Flash speed." It is live in the Gemini app and Search AI Mode, supports 1M token context and tool calling, and is priced at $0.50/M input tokens and $3.00/M output tokens.

> Links: Sundar Pichai tweet | Google blog | Google DeepMind announcement

xAI Launches Grok Voice Agent API for Real-Time Voice Interaction

The Grok Voice Agent API provides speech-to-speech interaction with tool calling, web/RAG search, and 100+ languages. It scores 92.3% on Big Bench Audio reasoning, priced at $0.05/minute.

> Links: xAI tweet | Benchmark analysis

Microsoft Releases TRELLIS 2-4B: Open-Source Image-to-3D Model

Microsoft released TRELLIS 2-4B, a 4-billion-parameter model using Flow-Matching Transformers and a sparse voxel 3D VAE to convert a single image into 3D assets. It is open-sourced on Hugging Face with a demo available.

> Links: Reddit post | Hugging Face page | Official blog

Apple Introduces SHARP: Photorealistic 3D Gaussians from a Single Image

Apple's SHARP generates photorealistic 3D Gaussian representations from a single image in seconds. It requires CUDA GPUs and the code is open-sourced on GitHub.

> Links: Reddit post | GitHub repo | arXiv paper

QwenLong-L1.5 Released: Long-Context Reasoning with 4M Tokens

QwenLong-L1.5 achieves SOTA long-context reasoning with support for 4 million tokens, using data synthesis, reinforcement learning, and memory management techniques. Open-sourced on Hugging Face.

> Links: Reddit post | Hugging Face page

---

Model Performance and Benchmarks

  • Gemini 3 Flash beats several mainstream models: It outperforms Gemini 3 Pro on ARC-AGI-2 and SWE-bench Verified, enters the Top 5 on LMArena and Vision Arena, and approaches GPT-5.2 on some metrics. (Benchmarks)
  • Mistral beats Gemini 3 Pro on ARC-AGI-2 with fewer parameters: Users speculate training methods forced better generalization rather than memorization. (ARC-AGI2)
  • QwenLong-L1.5 excels on long-context tasks: Users report it outperforms standard Qwen and Nemotron Nano on long-context information extraction, with attention needed on query templates.
  • ---

    Cost and Pricing

  • Gemini 3 Flash pricing: $0.50/M input tokens, $3.00/M output tokens — 75% cheaper than Pro, accessible via Google AI Studio and Vertex AI. (Pricing page)
  • Claude Opus API costs draw complaints: A Perplexity user reported ~$1.2 for 29K tokens; discussion on adding it to Perplexity Max. (Anthropic pricing)
  • OpenRouter timeout errors hit production: Multiple users reported timeouts (cURL 28) on the /completions endpoint, affecting production software, including some spending $6,000/month. (Status page)
  • ---

    Tools and Integrations

  • Gemini 3 Flash enhanced tool calling: Demos support 100+ tools; integrated into Cursor, VS Code, and Ollama Cloud.
  • Unsloth launches CLI tool: Installable in Python environments, enabling script-based training as an alternative to Jupyter notebooks. (GitHub)
  • Qdrant releases Snappy: An open-source multimodal PDF search pipeline using ColPali patch-level embeddings and multi-vector search, with production deployment guidance. (GitHub)
  • ---

    Multimodal and 3D Generation

  • Tencent releases Hunyuan HY World 1.5: Real-time interactive 3D world generation using Reconstituted Context Memory and Dual Action Representation, supporting first/third-person views. (Paper)
  • Runway Gen-4.5 emphasizes physically realistic motion: Kling 2.6 adds motion and voice control; TurboDiffusion claims 100–205× video generation speedup.
  • LangSmith launches observability tooling: OpenTelemetry tracing, pairwise preference queues, and automated evaluation, supporting large-scale agent deployments at Vodafone and Fastweb. (Case study)
  • ---

    User Experience and Feedback

  • GPT-5.2 criticism: Users cite "blatant hallucinations" and "stiff responses," with some switching to Gemini 3 Flash.
  • ChatGPT for bipolar disorder support: A user with bipolar II reports ChatGPT 5.1 helped manage obsessive thoughts and hypomanic episodes, providing non-judgmental support they found more effective than 5 years of traditional therapy.
  • Gemini 3 Flash feedback: Users praise tool calling and IDE integrations like Cursor, noting it beats Pro on benchmarks like SWE-bench.
  • ---

    Infrastructure and Hardware

  • RTX PRO 5000 Blackwell specs leaked: GB202 chip with 110 SMs enabled, 300W TDP, FP8/FP16 MMA with FP32 accumulation, and memory bandwidth at 3/4 of RTX 5090. (Datasheet)
  • GPU rental platform experiences vary: Network bandwidth differs widely across platforms like vast.ai; users recommend local debugging and setup scripts to reduce waste.
  • NeoCloudX offers budget GPU rentals: A100 at ~$0.4/hour and V100 at ~$0.15/hour by aggregating idle datacenter GPUs. (Website)
---

📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant

Tags

#ai-news#gemini-3-flash#grok#trellis#image-to-3d#long-context#gpu-rental#ai-daily

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169218