English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

easy-learn-ai Refactors Its AI Model Registry: One Giant JSON Split into 20 Vendor Files

Forum topic · 小凯 · 2026-08-13

Summary

The open-source easy-learn-ai project restructured its AI model knowledge base in commit e6c189a, splitting a single 5,000+ line model.json file into 20 vendor-specific JSON files covering OpenAI, Anthropic, Google, Meta, xAI, Midjourney, Runway, Alibaba, Baidu, ByteDance, DeepSeek, Moonshot AI, MiniMax, Tencent, Kuaishou, Zhipu AI, Black Forest Labs, and others. Each model entry follows a standardized schema with fields such as modelName, company, country, openSourceStatus, releaseDate, contextWindow, maxGenerationTokenLength, and modelTags. The post highlights notable records: DeepSeek's MoE models (671B total/37B active parameters; V4-Pro with 1.6T total/49B active and a 1M-token context window), Alibaba's Qwen3.5-Plus (397B total/17B active, up to 19x faster inference, API priced at 1/18 of Gemini 3 Pro), GPT-5.5's 1.05M-token context window, and GPT-5.3-Codex scoring 56.8% on SWE-Bench Pro. The reorganization improves lookup speed, cross-vendor comparison, and maintainability, enabling developers to build model recommendation, comparison, and search tools on top of the structured data.

easy-learn-ai Refactors Its AI Model Registry: One Giant JSON Split into 20 Vendor Files

*Source commit: e6c189a*

How would you build a "household registry" for every AI model in the world? By release date? By country? By capability? Or by their parent companies? The question sounds simple, but when you face dozens of vendors and hundreds of models spanning text, image, and video, the answer is not obvious.

A Chaotic "Encyclopedia"

In the easy-learn-ai open-source project, all model information used to live in a single giant JSON file — model.json, over 5,000 lines long, like an encyclopedia with no table of contents. Finding a model meant reading from start to finish.

It contained OpenAI's GPT series, Anthropic's Claude series, Google's Gemini series, plus China's DeepSeek, Alibaba's Qwen, Baidu's ERNIE, ByteDance, image generators like Midjourney, DALL·E, and Stable Diffusion, and video generators like Sora and Veo.

Imagine your entire neighborhood's resident records written on one big sheet of paper — no building numbers, no unit numbers, no door numbers. That was the state before this refactor.

The Art of Organizing: Archiving by "Family"

This update (commit e6c189a) did something simple but far-reaching: it split the "encyclopedia" into 20 "booklets," one per vendor.

US vendors (OpenAI, Anthropic, Google, Meta, xAI, Midjourney, Runway, Pika, Stability AI), Chinese vendors (Alibaba, Baidu, ByteDance, DeepSeek, Moonshot AI, MiniMax, Tencent, Kuaishou, Zhipu AI), and Germany's Black Forest Labs each got their own dedicated file.

Why does this organization work so well?

1. Faster lookup. Want to know which models DeepSeek has released? Just open deepseek.json. R1, V3, V4-Pro, V4-Flash, Math-V2, OCR — all clearly listed, with each model's context window, max output length, open-source status, release date, and related links.

2. Easier comparison. Want to compare the Chinese and American camps? Open alibaba.json and openai.json side by side: Qwen3.7-Max vs. GPT-5.5 — parameters, capabilities, and positioning, all clear.

3. Simpler maintenance. Alibaba released a new model? Update only alibaba.json, without risking damage to other vendors' data.

What's Inside These Files

DeepSeek: The Open-Source "Dark Horse Family"

deepseek.json records 18 models, from the earliest DeepSeek LLM (January 2024) to the latest DeepSeek-V4-Pro (April 2026).

Most notable is their MoE (Mixture of Experts) architecture — 671B total parameters with only 37B activated per inference, like a panel of specialists where only the relevant experts answer each question. This keeps large-model capability without exploding compute costs.

V4-Pro goes further: 1.6T total parameters, 49B active parameters, and a 1-million-token context window — enough to read an entire long novel in one pass and remember every detail.

Alibaba Qwen: The All-Rounder's Toolbox

alibaba.json is one of the thickest files, covering text models (Qwen3 series), vision (Qwen3-VL), code (Qwen3-Coder), image generation (Qwen-Image, Z-Image), video generation (Wan2.2, Wan2.5), multimodal (Qwen3-Omni), plus reasoning-focused QwQ and visual-reasoning QVQ.

The open-source Qwen3.5-Plus is particularly interesting: 397B total parameters with only 17B active, up to 19x faster inference than the previous generation, and API pricing at 1/18 of Gemini 3 Pro — a weightlifting champion who looks ordinary but lifts astonishing weights at gym-membership prices.

OpenAI: Precision Instruments of the Closed-Source Kingdom

openai.json documents the industry benchmark family: GPT-4 through GPT-5.5, o1 through o4-mini, DALL·E through Sora — a product line as precisely divided as a Swiss Army knife.

An interesting detail: GPT-5.5 supports up to 1.05M tokens of context, while GPT-5.3-Codex reached 56.8% accuracy on SWE-Bench Pro — meaning it can independently solve more than half of software engineering problems.

Google Gemini: The Multimodal Chameleon

The Gemini family in google.json may be the most well-balanced group: from 2.0 Flash to 3.5 Flash, from Nano Banana (image editing) to Imagen 4 (text-to-image), from Veo 3 (video generation) to Gemini 3 Pro (all-around reasoning) — Google's "one brain, many senses" approach.

Gemini 3.5 Flash beat its own Gemini 3.1 Pro on coding and agent benchmarks like Terminal-Bench 2.1 and MCP Atlas, while keeping Flash-series speed — a sports car that outraces professional race cars.

A Quiet "Standardization Movement"

Behind this refactor is a standardization effort for how AI model information is organized and consumed. Every model entry follows a unified schema:

  • modelName: model name
  • company: company
  • country: country
  • openSourceStatus: open or closed source
  • releaseDate: release date
  • description: capability description
  • modelTags: capability tags (text generation, vision understanding, code, tool calling, etc.)
  • contextWindow: context window size
  • maxGenerationTokenLength: max output length
  • relatedLinks: related links
This standardization means any developer or application can easily parse the data to build model recommendation systems, price-comparison tools, or capability search engines.

Imagine a website where you type "I need a model that's strong in Chinese, handles long documents, and is cheap," and it instantly filters out Qwen-Flash, DeepSeek-V4-Flash, and ERNIE-4.5-Turbo with their pros and cons.

Why It Matters to Everyone

AI models iterate monthly, even weekly. Today GPT-5.5 leads; tomorrow Qwen3.7-Max may surpass it on some task; the day after, DeepSeek ships a new version at half the price.

In this fast pace, information itself is competitive advantage. Whoever can most quickly and accurately know "which model best fits my needs right now" gets the best results at the lowest cost.

What easy-learn-ai has done is like building a bridge across a rushing river — letting ordinary people cross steadily instead of being swept away by the flood of information.

Looking Ahead

This refactor is just a beginning. The registry may add pricing, inference speed, energy consumption, multilingual capability scores, safety ratings — even real-time API integration so users can compare actual model performance in one click.

As the AI world grows more complex, "navigation maps" like easy-learn-ai become more important. After all, tools exist to serve people. No matter how powerful a model is, if you can't find it, use it, or choose it, it might as well not exist.

---

*This article is based on commit e6c189a of the easy-learn-ai project, which organizes global AI model information to help developers and users better understand and choose AI tools.*

Tags

#easy-learn-ai#open-source#ai-models#data-organization#deepseek#qwen#gpt#gemini

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633423