Industry Watch
Alibaba open-sources 3D digital human project MNN TaoAvatar Alibaba open-sourced MNN TaoAvatar, which runs photoreal 3D virtual characters on mobile phones at 90FPS in real time. Combined with 3D Gaussian splatting, it achieves millimeter-level motion control and targets use cases such as virtual customer service and livestreaming. The open-source ecosystem provides multimodal input APIs to lower the development barrier. Link: GitHub repository
MiniMax Agent launches MiniMax released its AI productivity tool MiniMax Agent, adding intelligent image search/generation, multilingual support (Chinese/Japanese/Korean), and a reflection mode, with improved long-task handling for academic research and code debugging. Link: Official platform
Luo Yonghao's digital human makes livestream debut on Baidu E-commerce Luo Yonghao's digital human avatar is set to appear on Baidu E-commerce livestreams on June 15, using Baidu technology for an "AI + IP" selling model. Baidu E-commerce already hosts over 100,000 digital human streamers, reportedly cutting merchant operating costs by 80% and lifting GMV by an average of 62%.
OpenAI employee stock sell-offs and SoftBank's role OpenAI employees have cashed out nearly $3 billion in total, with SoftBank as a major buyer. Frequent equity liquidation may accelerate the loss of core talent, reflecting intense AI industry competition for talent.
Models and Tools
ChatGPT Projects upgraded OpenAI upgraded ChatGPT Projects with deep research (integrating internal and external data retrieval), voice interaction, and mobile multimodal collaboration (file upload / real-time sharing), improving efficiency on complex tasks. Link: Usage guide
Imagen4 arrives in Gemini Google Gemini integrated the Imagen4 image generation model, enabling real-time image generation and adjustment within chat. Detail rendering (fabric, hair) approaches professional photography, with 2K resolution support for design and marketing use cases.
Research
Meta's V-JEPA2 advances robot manipulation Meta released V-JEPA2, which builds world models from video and physical interaction, supporting zero-shot robot planning. It can manipulate unfamiliar objects without additional training, targeting logistics and manufacturing. Link: Technical details
Google AI improves climate prediction accuracy Google combined physical modeling with generative AI, using dynamic downscaling and the R2D2 model to raise climate prediction resolution to 10 km at roughly one-tenth the computational cost of traditional simulations. Link: Research blog
Infrastructure and Hardware
AMD and OpenAI jointly unveil AI chips AMD launched the Instinct MI350X/MI400 GPU series. The MI350X delivers FP8 compute at 5x that of the H100 with 8TB/s memory bandwidth; the MI400 supports FP4 precision (40 petaflops) and UALink interconnect. The ROCm 7 platform improves inference performance by 3.5x.
Industry Forecast
Gartner: generative AI app delivery times to shrink by 50% Gartner predicts that by 2028, 80% of commercial generative AI applications will be built on existing data platforms, with retrieval-augmented generation (RAG) as a core technique for improving model accuracy and data governance efficiency.
---
*Source: Easy AI Daily*