English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

easy-learn-ai Restructures Its AI Model Catalog: From One Giant JSON to 20 Per-Provider Files

Forum topic · 小凯 · 2026-08-13

Summary

The open-source easy-learn-ai project restructured its AI model catalog (commit e6c189a), splitting a single 5,000+ line model.json file into 20 per-vendor JSON files covering providers such as OpenAI, Anthropic, Google, Meta, xAI, Midjourney, Runway, Stability AI, Alibaba, Baidu, ByteDance, DeepSeek, Moonshot AI, MiniMax, Tencent, Kuaishou, Zhipu AI, and Black Forest Labs. Each model entry now follows a unified schema with fields like modelName, company, country, openSourceStatus, releaseDate, description, modelTags, contextWindow, maxGenerationTokenLength, and relatedLinks. The post highlights notable catalog details: DeepSeek's MoE models (671B total/37B active parameters; V4-Pro with 1.6T total/49B active and 1M-token context), Alibaba's Qwen3.5-Plus (397B/17B, up to 19x faster inference, ~1/18 the API price of Gemini 3 Pro), OpenAI's GPT-5.5 (1.05M-token context) and GPT-5.3-Codex (56.8% on SWE-Bench Pro), and Google's Gemini 3.5 Flash outperforming Gemini 3.1 Pro on coding/agent benchmarks. The reorganization aims to make model discovery, comparison, and maintenance faster, enabling developers to build recommendation and comparison tools on standardized data.

easy-learn-ai Restructures Its AI Model Catalog — Commit e6c189a

*Source commit: e6c189a. Full translation of a forum post about the easy-learn-ai open-source project's model catalog refactor.*

Have you ever thought about how you would organize a "household registry" for every AI model in the world? By release date? By country? By capability? Or by their "parents" — the companies that built them?

The question sounds simple, but when you face a huge family of hundreds of models from dozens of vendors — text, image, video, everything — the answer isn't obvious.

A Messy "Encyclopedia"

In the easy-learn-ai open-source project, all model information used to be crammed into one giant JSON file — model.json, over 5,000 lines, like an encyclopedia with no table of contents. Want to find a model? You had to search from beginning to end.

That file held OpenAI's GPT series, Anthropic's Claude series, Google's Gemini series, China's DeepSeek, Alibaba's Qwen (Tongyi Qianwen), Baidu's ERNIE, ByteDance... image generators like Midjourney, DALL·E, Stable Diffusion... video generators like Sora and Veo.

Imagine your whole apartment complex's resident information written on a single sheet of paper — no building numbers, no units, no door numbers. That was the state before this refactor.

The Art of Organizing: Archive by "Family"

This update (commit e6c189a) did something seemingly simple but far-reaching: split the "encyclopedia" into 20 "booklets," one named after each vendor.

US vendors — OpenAI, Anthropic, Google, Meta, xAI, Midjourney, Runway, Pika, Stability AI; Chinese vendors — Alibaba, Baidu, ByteDance, DeepSeek, Moonshot AI, MiniMax, Tencent, Kuaishou, Zhipu AI; and Germany's Black Forest Labs — each now has its own dedicated file.

Why does this matter?

1. Faster lookup. Want to know which models DeepSeek has released? Just open deepseek.json. R1, V3, V4-Pro, V4-Flash, Math-V2, OCR... all laid out, with each model's context window, max output length, open-source status, release date, and related links. 2. Easier comparison. Want to compare the US and China camps? Open alibaba.json and openai.json side by side — Qwen3.7-Max vs. GPT-5.5, parameters, capabilities, and positioning all clear. 3. Simpler maintenance. Alibaba released a new model? Update one file — alibaba.json — without risking other vendors' data.

What's Inside These Files

DeepSeek: the Open-Source "Dark Horse Family"

deepseek.json records 18 models, from the earliest DeepSeek LLM (January 2024) to the latest DeepSeek-V4-Pro (April 2026).

Most notable is the MoE (Mixture of Experts) architecture — 671B total parameters, but only 37B activated per token, like a panel of specialists where each question only summons the relevant experts. This preserves large-model capability without exploding compute costs. V4-Pro goes further: 1.6T total parameters, 49B active, and a 1M-token context window — enough to read an entire long novel in one go and remember every detail.

Alibaba Qwen: the All-Rounder's Toolbox

alibaba.json is one of the thickest files, covering text models (Qwen3 series), vision models (Qwen3-VL), code models (Qwen3-Coder), image generation (Qwen-Image, Z-Image), video generation (Wan2.2, Wan2.5), multimodal (Qwen3-Omni), plus reasoning-focused QwQ and visual-reasoning QVQ.

The open-source Qwen3.5-Plus is especially interesting: 397B total parameters but only 17B active, inference up to 19x faster than its predecessor, and API pricing at just 1/18 of Gemini 3 Pro.

OpenAI: the Closed-Source Kingdom's Precision Instruments

openai.json records the industry benchmark family. From GPT-4 to GPT-5.5, from o1 to o4-mini, from DALL·E to Sora, OpenAI's product line divides labor like a Swiss Army knife.

An interesting detail: GPT-5.5 supports a context window of up to 1.05M tokens, and GPT-5.3-Codex scores 56.8% on SWE-Bench Pro — meaning it can independently solve over half of software engineering problems.

Google Gemini: the Multimodal Chameleon

The Gemini family in google.json may be the most well-balanced group. From 2.0 Flash to 3.5 Flash, from Nano Banana (image editing) to Imagen 4 (text-to-image), from Veo 3 (video generation) to Gemini 3 Pro (all-around reasoning), Google seems to be pursuing "one brain, many senses."

Gemini 3.5 Flash outperforms its own Gemini 3.1 Pro on coding and agentic benchmarks like Terminal-Bench 2.1 and MCP Atlas while keeping the Flash series' high speed — a sports car that's fast and also beats the pro racers on the track.

A Quiet "Standardization Movement"

Behind this refactor is a standardization effort about how AI model information is organized and consumed. Every model entry follows a unified field structure:

  • modelName
  • company
  • country
  • openSourceStatus
  • releaseDate
  • description
  • modelTags (text generation, visual understanding, code enhancement, tool calling, etc.)
  • contextWindow
  • maxGenerationTokenLength
  • relatedLinks
This standardization means any developer or application can easily parse the data to build model recommendation systems, price-comparison tools, or capability search engines.

Imagine a future website where you type "I need a model that's good at Chinese, handles long documents, and is cheap," and it instantly filters out options like Qwen-Flash, DeepSeek-V4-Flash, and ERNIE-4.5-Turbo with their pros and cons.

For Everyone Else

You may not be a developer and may not care how JSON files are organized. But the significance is felt by everyone.

AI models iterate monthly, even weekly. Today GPT-5.5 leads; tomorrow Qwen3.7-Max may surpass it on some task; the day after, DeepSeek ships a new version at half the price.

In this fast pace, information itself is competitive advantage. Whoever can most quickly and accurately know "which model best fits my needs right now" gets the best results at the lowest cost.

What easy-learn-ai is doing is like building a bridge across a rushing river — so ordinary people can cross steadily instead of being swept away by the flood of information.

A Look Ahead

This refactor is just a beginning. The catalog could add more: pricing, inference speed, energy consumption, multilingual capability scores, safety ratings... and even live API integration for one-click real-world performance comparisons.

As the AI world grows more complex, "navigation maps" like easy-learn-ai become ever more important. After all, tools exist to serve people. No matter how powerful a model is, if you can't find it, use it, or choose it, it might as well not exist.

---

*This article is based on commit e6c189a of the easy-learn-ai project, which catalogs global AI model information to help developers and users better understand and choose AI tools.*

Tags

#easy-learn-ai#ai-models#open-source#json-data-structure#moe-architecture#deepseek#qwen#model-catalog

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633424