← All projects
Projects 05/Private · Python, LLM agents

ace-trading-lab

An experiment in running two LLM traders, GPT and Grok, side by side on identical rules, with a deterministic risk cage between every decision and every order. Paper and testnet accounts only. The point is the architecture, the evaluation and the controls, not returns.

Private repository

01What it is

Two language models trade the same universe of crypto pairs on separate virtual books: GPT through the Codex CLI and Grok through the Grok Build CLI, both on existing subscriptions. Each cycle they see the same market snapshot and answer with a structured decision. It is an engineering experiment about how far you can trust a model that is allowed to act, and it is not trading advice.

02What I built

A Python and FastAPI service with a scheduler, a market and news collector, an orchestrator, pluggable venue adapters for the Binance spot testnet and an Interactive Brokers paper account, and a SQLite store for decisions, orders, fills and audit events. The risk guard is plain code: position and exposure caps, a daily order limit, a minimum gap between orders, long-only, a venue allowlist and a drawdown kill switch. Every rejection is stored with the rule that fired. A React dashboard on Cloudflare Pages shows the equity curves, the reasoning behind each decision and the risk events.

03How it runs

A cycle collects prices, news headlines and a sentiment index, asks both traders the same question and compares their answers, with a combine mode that is implemented but switched off. A reply that does not parse, a timeout or a rate limit becomes a hold plus an audit event, never a retry loop. Settings can be tuned from the dashboard, but the API validates every value and the dashboard is never trusted. The backend deploys to a VPS and the frontend to Cloudflare Pages from GitHub Actions, only after the tests pass, and the API is reached through a Cloudflare tunnel behind Cloudflare Access.

04What I learned

The traders never grade themselves: a separate auditor reads the database and scores expectancy and fees, not raw profit and loss. A bug taught me the most. A resting limit order filled on the exchange but was never settled in my database, so one trader looked like it had lost a large share of its capital. I added a reconciliation step that compares the exchange fills with my own records each cycle. I also learned to make a throttled trader distinguishable from a cautious one, because a quiet agent can mean two very different things.

How it is built

  1. Market context

    Public data in

  2. Decision cycle

    Runs on a schedule

  3. LLM traders

    Same rules, same input

  4. Risk cage

    Deterministic code

  5. Evidence and delivery

    Decisions, monitoring, delivery

Select a part of the diagram to see what it does. Each color is one flow through the system.

© 2026 Benoit Ardiet · Quito, Ecuador · UTC−5GitHubLinkedInNo tracking, no cookies.