Easy AI Daily | December 18, 2025 AI Industry News
Model Releases and Updates
Google Releases Gemini 3 Flash: Pro-Level Reasoning at 1/4 the Cost
Gemini 3 Flash is positioned as "Pro-level reasoning at Flash speed." It is live in the Gemini app and Search AI Mode, supports 1M token context and tool calling, and is priced at $0.50/M input tokens and $3.00/M output tokens.> Links: Sundar Pichai tweet | Google blog | Google DeepMind announcement
xAI Launches Grok Voice Agent API for Real-Time Voice Interaction
The Grok Voice Agent API provides speech-to-speech interaction with tool calling, web/RAG search, and 100+ languages. It scores 92.3% on Big Bench Audio reasoning, priced at $0.05/minute.> Links: xAI tweet | Benchmark analysis
Microsoft Releases TRELLIS 2-4B: Open-Source Image-to-3D Model
Microsoft released TRELLIS 2-4B, a 4-billion-parameter model using Flow-Matching Transformers and a sparse voxel 3D VAE to convert a single image into 3D assets. It is open-sourced on Hugging Face with a demo available.> Links: Reddit post | Hugging Face page | Official blog
Apple Introduces SHARP: Photorealistic 3D Gaussians from a Single Image
Apple's SHARP generates photorealistic 3D Gaussian representations from a single image in seconds. It requires CUDA GPUs and the code is open-sourced on GitHub.> Links: Reddit post | GitHub repo | arXiv paper
QwenLong-L1.5 Released: Long-Context Reasoning with 4M Tokens
QwenLong-L1.5 achieves SOTA long-context reasoning with support for 4 million tokens, using data synthesis, reinforcement learning, and memory management techniques. Open-sourced on Hugging Face.> Links: Reddit post | Hugging Face page
---
Model Performance and Benchmarks
- Gemini 3 Flash beats several mainstream models: It outperforms Gemini 3 Pro on ARC-AGI-2 and SWE-bench Verified, enters the Top 5 on LMArena and Vision Arena, and approaches GPT-5.2 on some metrics. (Benchmarks)
- Mistral beats Gemini 3 Pro on ARC-AGI-2 with fewer parameters: Users speculate training methods forced better generalization rather than memorization. (ARC-AGI2)
- QwenLong-L1.5 excels on long-context tasks: Users report it outperforms standard Qwen and Nemotron Nano on long-context information extraction, with attention needed on query templates.
- Gemini 3 Flash pricing: $0.50/M input tokens, $3.00/M output tokens — 75% cheaper than Pro, accessible via Google AI Studio and Vertex AI. (Pricing page)
- Claude Opus API costs draw complaints: A Perplexity user reported ~$1.2 for 29K tokens; discussion on adding it to Perplexity Max. (Anthropic pricing)
- OpenRouter timeout errors hit production: Multiple users reported timeouts (cURL 28) on the /completions endpoint, affecting production software, including some spending $6,000/month. (Status page)
- Gemini 3 Flash enhanced tool calling: Demos support 100+ tools; integrated into Cursor, VS Code, and Ollama Cloud.
- Unsloth launches CLI tool: Installable in Python environments, enabling script-based training as an alternative to Jupyter notebooks. (GitHub)
- Qdrant releases Snappy: An open-source multimodal PDF search pipeline using ColPali patch-level embeddings and multi-vector search, with production deployment guidance. (GitHub)
- Tencent releases Hunyuan HY World 1.5: Real-time interactive 3D world generation using Reconstituted Context Memory and Dual Action Representation, supporting first/third-person views. (Paper)
- Runway Gen-4.5 emphasizes physically realistic motion: Kling 2.6 adds motion and voice control; TurboDiffusion claims 100–205× video generation speedup.
- LangSmith launches observability tooling: OpenTelemetry tracing, pairwise preference queues, and automated evaluation, supporting large-scale agent deployments at Vodafone and Fastweb. (Case study)
- GPT-5.2 criticism: Users cite "blatant hallucinations" and "stiff responses," with some switching to Gemini 3 Flash.
- ChatGPT for bipolar disorder support: A user with bipolar II reports ChatGPT 5.1 helped manage obsessive thoughts and hypomanic episodes, providing non-judgmental support they found more effective than 5 years of traditional therapy.
- Gemini 3 Flash feedback: Users praise tool calling and IDE integrations like Cursor, noting it beats Pro on benchmarks like SWE-bench.
- RTX PRO 5000 Blackwell specs leaked: GB202 chip with 110 SMs enabled, 300W TDP, FP8/FP16 MMA with FP32 accumulation, and memory bandwidth at 3/4 of RTX 5090. (Datasheet)
- GPU rental platform experiences vary: Network bandwidth differs widely across platforms like vast.ai; users recommend local debugging and setup scripts to reduce waste.
- NeoCloudX offers budget GPU rentals: A100 at ~$0.4/hour and V100 at ~$0.15/hour by aggregating idle datacenter GPUs. (Website)
---
Cost and Pricing
---
Tools and Integrations
---
Multimodal and 3D Generation
---
User Experience and Feedback
---
Infrastructure and Hardware
📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant