TradingAgents: How LLM Agents Replicate a Wall Street Trading Firm's Division of Labor
The Scenario
Imagine walking into a Wall Street trading firm: fundamental analysts on the left, sentiment analysts on the right, a news analyst watching Bloomberg terminals, a technical analyst staring at candlestick charts. In the back room, bull and bear research teams debate across the table. A trader synthesizes everyone's input and places orders, a risk team monitors volatility and liquidity, and finally a portfolio manager signs off.
TauricResearch's TradingAgents does exactly this—it ports that division of labor into LLMs. Each role is played by an LLM agent; agents collaborate through structured dialogue and ultimately produce a trading decision. It's not a single model making a snap judgment, but an entire LLM team running a process.
Architecture: A Trading Firm in LLM Form
The core design is role specialization + debate. The framework has four layers:
Analyst team — four agents, each with its own domain:
- Fundamentals analyst: reads earnings reports and performance metrics to find intrinsic value and red flags
- Sentiment analyst: aggregates headlines, StockTwits, and Reddit posts into a single sentiment reading
- News analyst: monitors global news and macro indicators, interpreting event impact
- Technical analyst: detects trading patterns with indicators like MACD and RSI
- v0.2.0: multi-provider support (GPT-5.x, Gemini 3.x, Claude 4.x, Grok 4.x)
- v0.2.3: multilingual support, GPT-5.4, backtest date fidelity
- v0.2.4: structured-output agents, LangGraph checkpoint resume, persistent decision logs, Docker support
- v0.2.5: grounded sentiment analyst, Qwen/GLM/MiniMax dual-region support, Ollama local models
- v0.3.0: validated data access contracts, expanded provider registry (NVIDIA, Kimi, Groq, Mistral, Bedrock), FRED and Polymarket data sources
- v0.3.1: Alpha Vantage look-ahead filtering, graph-router crash safety, graph-shape-aware checkpoint resume, Claude Sonnet 5 / Fable 5 support
- Octopus-style DNA pretraining + RNA inference-time computation—different layers for different problems
- SOPHIA's differential flow directions—different states use different exit directions
- Euclid-MCP's LLM + Prolog—LLM as poet, Prolog as accountant
- Rebucca's small model + large model review—small model pre-screens, large model verifies
- GitHub: https://github.com/TauricResearch/TradingAgents
- Paper: https://arxiv.org/abs/2412.20138 (267 citations)
- Latest version: v0.3.1
- Supported providers: OpenAI / Anthropic / Google / xAI / DeepSeek / Qwen / GLM / Mistral / Groq / NVIDIA / Bedrock / Ollama
- License: MIT
Researcher team — one bull and one bear agent critically evaluate the analysts' insights through structured debate, balancing upside and risk. This step is key: they argue before any order is placed.
Trader agent — synthesizes analyst reports and the researcher debate to decide trade timing and size.
Risk team + portfolio manager — the risk team continuously evaluates volatility, liquidity, and other factors, then issues an assessment. The portfolio manager approves or rejects the proposal. Approved orders go to a simulated exchange.
The most notable aspect: it doesn't ask one LLM to do everything—it decomposes trading decisions into specialized roles, each agent handling only its own part. This mirrors real trading firm structures by deliberate design, not coincidence.
Academic Backing: 267 Citations
TradingAgents is grounded in a research paper, arXiv:2412.20138 (December 2024), which has been cited 267 times—a solid count in AI for Finance.
The paper's core contribution is showing that multi-agent debate produces more stable decisions than single-agent reasoning. The bull/bear debate isn't theater; adversarial discussion surfaces blind spots, matching how human investment committees work.
Engineering: From Paper to Open-Source Framework
Released in January (first release) and now at v0.3.1, iteration has been dense:
Several engineering decisions stand out:
1. No provider lock-in: supports OpenAI, Anthropic, Google, xAI, DeepSeek, Qwen, GLM, Mistral, Groq, NVIDIA, Bedrock, and any OpenAI-compatible endpoint. You can use different models for different roles—e.g., GPT-5.5 for fundamentals, Claude 4.6 for debate, local Ollama for sentiment. 2. LangGraph checkpoint resume: the pipeline (analysts → researchers → trader → risk) can be long; checkpoints let you recover from mid-run crashes. 3. Look-ahead filtering: filters future information during backtests to prevent data leakage inflating returns—a classic quant pitfall, specifically fixed in v0.3.1. 4. Structured outputs: Research Manager, Trader, and Portfolio Manager emit structured output rather than free text, so downstream agents can reliably parse upstream results.
The Pattern: Specialization Beats Monoliths
TradingAgents exemplifies a recurring engineering principle: division of labor beats uniformity.
TradingAgents is this principle applied to finance: rather than one LLM doing everything, specialists do specialist work, integrated via structured dialogue.
A hidden advantage is debuggability. When a trade goes wrong, you can locate which stage failed—did the sentiment analyst misread Reddit? Was the bull researcher's debate too shallow? Did risk miss a liquidity problem? A single-LLM black box can't offer this granularity of attribution.
Real-World Limitations
TradingAgents explicitly states it is for research only and not investment advice. Trading performance depends on many factors: backbone model choice, temperature, trading horizon, data quality, and non-determinism. The framework currently connects to a simulated exchange, not real markets.
That's a reasonable limit—LLM agents aren't yet reliable enough to manage real money. But as a research framework, it offers a reproducible, debuggable, comparable experimental platform for multi-agent trading.
Data
Why It Matters
TradingAgents grounds the abstract idea of "multi-agent collaboration" in a concrete domain—financial trading. Its value isn't just trading itself, but demonstrating how to map human organizational structures onto LLM agent architectures. The analysts → researchers → trader → risk → portfolio manager chain corresponds one-to-one with a real trading firm.
This "organizational isomorphism" design transfers to other multi-role domains—medical diagnosis (symptom analysis → differential diagnosis → treatment decisions → risk review), legal analysis (fact gathering → statute matching → debate → review), even code review (static analysis → security audit → performance evaluation → architecture review).
As LLM agents move from solo operation to teamwork, TradingAgents offers a paradigm well worth studying.