Easy AI Daily | January 17, 2026
Key points
Products & Applications
- OpenAI launches ChatGPT Go ($8/month) — 10x more messages than Free, file uploads, image generation, longer memory/context, and unlimited GPT-5.2 instant. Ads will be tested on Free and Go tiers; Plus/Pro/Business/Enterprise remain ad-free. (Announcement | Ad principles)
- Claude Cowork opens to Pro users — Still a research preview; users report organizing 400+ files consumed 97% of session quota, sparking complaints about tight limits. (Reddit discussion)
- Gemini 3 Pro performance complaints — Pro users report degraded long-project performance and suspect a silently reduced context window; some are migrating to GPT-5.2 Thinking or Claude. (Reddit feedback)
- Perplexity Pro's 100/day advanced-query cap — Heavy users exhaust quotas within hours and are considering switching to token-billed alternatives.
- AI IDE/CLI costs under fire — Reported bills: Cursor Ultra burning 20% of quota per orchestrator run, Qoder near $400/month, Gemini CLI ~$120/day at 10M tokens. Users want clearer usage dashboards and hybrid model routing.
- LMArena adds PDF chat testing and updated image leaderboards; Hawk Ultra (Movement Labs) is hyped for generating 9k–20k lines of code per prompt as a possible "Opus killer."
- Sam Altman teases "Very fast Codex" and improved ChatGPT memory; developers discuss workflows shifting to fast models with human-in-the-loop steering. (Tweet)
- Codex CLI supports open-weight models via
codex --oss(Ollama), with 32K+ context recommended and mid-generation steering being tested. - SWE-rebench (Dec 2025): Claude Opus 4.5 leads at 63.3%, GPT-5.2 xhigh at 61.5%; Gemini 3 Flash Preview beats its own Pro; GLM-4.7 is the strongest open-source model. (swe-rebench.com)
- Unsloth claims 7–12x longer RL context (20K on 24GB VRAM, up to 380K on a 192GB B200) via data migration and new batching algorithms.
- Zhipu & Huawei release GLM-Image — trained fully on Ascend 910, supports 1024–2048 resolutions without extra training, strong Chinese text rendering, ~60% better tokens-per-joule than H200 (claimed), API ~0.1 RMB/image.
- VoxCPM (OpenBMB) — open-source token-free streaming TTS with LoRA fine-tuning, ~0.15 real-time factor on a single 4090.
- Translate Gemma now on Hugging Face and integrated into Ollama; OpenBMB AIR alignment framework reports +5.3 average points across 6 benchmarks using 14K curated samples.
- Human-in-the-loop validated again as a major reliability multiplier for agentic workflows.
- Jerry Liu (LlamaIndex): fixed chunking + vector DB RAG is dying for small document sets — direct file tools (ls/grep) beat pre-chunked embedding until scale demands a database.
- Claude and OpenRouter now support multiple parallel tool calls in a single request, cutting latency and cost.
- New orchestration tools emerging: SpecStory CLI, sled UI, OpenWork local computer agents; Claude Flow v3 claims 2.5x effective Claude Max capacity via WASM multi-agent swarms, though the community questions its benchmarks.
- "Inference explosion year": a widely shared essay argues prefill now dominates cost, context caching becomes standard, and scheduling/memory hierarchies need rework.
- SambaNova SN40L runs DeepSeek R1, beating NVIDIA clusters on high-concurrency throughput (~269 tok/s single-user peak).
- Epoch AI estimates ~30GW of installed AI datacenter power — roughly New York State's summer peak.
- Kernel engineering: NVIDIA CuTe/cuTile tiling nearing cuBLAS performance; AMD gfx942 multi-L2 coherence requires
buffer_inv sc1to avoid stale caches. - PCIe matters: a 3090 on Gen3 x1 drops inference from 120 to 90 tok/s;
sleep(2s)in benchmarks causes GPU downclocking artifacts. - GPU market: a working used A100 40GB found for $500; RTX 5060 Ti 16GB reportedly discontinued/reduced, raising prices.
- Mamba-2 rewrote its core scan as block-diagonal GEMM, lifting Tensor Core utilization from 10–20% to 60–70%; RetNet's abandonment shows the Transformer–hardware–resources lock-in.
- Multi-vector retrieval (ColBERT/ColPali-style): a 32M-parameter model with multi-vector can approach 8B-model retrieval quality.
- Information Gravity proposal (GPU MODE) models hallucination loops via excitation thresholds with a hysteresis firewall — currently more thought experiment than method.
- MMLU-Pro fixed: EleutherAI patched the dataset and lm-evaluation-harness; older MMLU-Pro scores may be skewed and should be re-run. (PR)
- OpenAI monetizes ~900M weekly users via ads + tiered subscriptions, seen as a pivot toward a hybrid ad/subscription model.
- Higgsfield AI raised $130M at a $1.3B valuation, claiming $200M annualized revenue in under 9 months.
- A tax-automation startup (Saket Kumar) raised a $3.5M seed from General Catalyst to make US personal tax filing free and one-click.
- Anthropic's 4th Economic Index introduces "economic primitives" (task complexity, education level, autonomy, success rate) to quantify AI's labor-market impact.
- OpenAI ad principles: answers won't be altered by advertisers, ads clearly labeled, conversations not shared with advertisers — but the community worries about long-term incentive drift.
- BASI Jailbreaking community catalogs daily jailbreak techniques (Gemini NSFW bypasses, Llama3 refusal reversal, OCR/cool-link filter evasion) that vendors rapidly patch.
- A zero-knowledge-proof AI moderation proposal would let platforms verify content was screened without exposing the content itself.
- AAAI 2026 will host a Machine Consciousness workshop (CIMC), with submissions due January 23, focusing on operational detection methods rather than philosophy.
Models & Capabilities
Agents & Tooling
Infrastructure & Hardware
Research & Methods
Industry & Business
Policy, Governance & Safety
📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant