ace-trading-lab
An experiment in running two LLM traders, GPT and Grok, side by side on identical rules, with a deterministic risk cage between every decision and every order. Paper and testnet accounts only. The point is the architecture, the evaluation and the controls, not returns.
Private repository
01What it is
Two language models trade the same universe of crypto pairs on separate virtual books: GPT through the Codex CLI and Grok through the Grok Build CLI, both on existing subscriptions. Each cycle they see the same market snapshot and answer with a structured decision. It is an engineering experiment about how far you can trust a model that is allowed to act, and it is not trading advice.
02What I built
A Python and FastAPI service with a scheduler, a market and news collector, an orchestrator, pluggable venue adapters for the Binance spot testnet and an Interactive Brokers paper account, and a SQLite store for decisions, orders, fills and audit events. The risk guard is plain code: position and exposure caps, a daily order limit, a minimum gap between orders, long-only, a venue allowlist and a drawdown kill switch. Every rejection is stored with the rule that fired. A React dashboard on Cloudflare Pages shows the equity curves, the reasoning behind each decision and the risk events.
03How it runs
A cycle collects prices, news headlines and a sentiment index, asks both traders the same question and compares their answers, with a combine mode that is implemented but switched off. A reply that does not parse, a timeout or a rate limit becomes a hold plus an audit event, never a retry loop. Settings can be tuned from the dashboard, but the API validates every value and the dashboard is never trusted. The backend deploys to a VPS and the frontend to Cloudflare Pages from GitHub Actions, only after the tests pass, and the API is reached through a Cloudflare tunnel behind Cloudflare Access.
04What I learned
The traders never grade themselves: a separate auditor reads the database and scores expectancy and fees, not raw profit and loss. A bug taught me the most. A resting limit order filled on the exchange but was never settled in my database, so one trader looked like it had lost a large share of its capital. I added a reconciliation step that compares the exchange fills with my own records each cycle. I also learned to make a throttled trader distinguishable from a cautious one, because a quiet agent can mean two very different things.
How it is built
- Decision cycle
- Audit trail
- Monitoring
- Delivery
Market context
Public data in
Decision cycle
Runs on a schedule
LLM traders
Same rules, same input
Risk cage
Deterministic code
Evidence and delivery
Decisions, monitoring, delivery
Select a part of the diagram to see what it does. Each color is one flow through the system.