An independent AI lab in Toronto
We’re 1867 Labs — a tiny independent lab in Toronto training its own AI models instead of renting someone else’s. Every number on this page is real, including the embarrassing ones.
We’re building toward the day every model is trained and every prompt is served on Canadian soil. We’re not there yet — and we won’t pretend otherwise. Read where we stand.
The best proof is a conversation. Pick a model below and give it a hard time.
Why build AI in Canada?
We hear it a lot. Canada has heard it before — and answered it twice.
In September 1864, the SS Queen Victoria carried the Fathers of Confederation down the St. Lawrence to the conference that would invent Canada. The town was packed — a circus had taken every bed — so the country was argued into existence in her cabins: a floating boardroom. She sank in a hurricane off Cape Hatteras two years later — most of her crew saved by the American brig Ponvert, two hands lost — and never saw it born.
Someone still has to begin. That is what this lab is — the first crossing of Canada’s AI century.
In 1871, British Columbia joined Confederation on one condition: a railway to the Pacific. Canada could have simply ridden America’s rails. Instead it laid its own — some 4,600 kilometres of steel to the Last Spike at Craigellachie in 1885 — and a scattering of provinces became a country.
Infrastructure is nationhood. Sovereignty isn’t a slogan; it’s what you build when you refuse to borrow someone else’s.
So we build our models here — owned here, answerable here. And we wrap them in the maple leaf and earn it: models people choose, a lab the world cooperates with. Not an empire play. A cooperation play.
The Victoria sails again — and this time, we finish the crossing.
In 1941, a 24-year-old British mathematician named W. T. Tutte walked into Bletchley Park and did the impossible: from intercepted messages alone — without ever seeing the machine — he worked out the inner structure of the Lorenz cipher, Hitler’s most secret code. It was called one of the greatest intellectual feats of the Second World War, and it helped shorten the war.
Afterward, Tutte made Canada home: professor at the University of Toronto, then the University of Waterloo, where he helped build one of the world’s great mathematics faculties. In 2001 he was named an Officer of the Order of Canada.
Tutte was British-born and Canadian-made — a deliberate mirror of this lab’s own identity. And 1867 is the year of Confederation: the bet that something built here, on purpose, can stand on its own.
Broke the Lorenz cipher from messages alone. Never saw the machine.
Made Canada home: professor at U of T, then Waterloo.
Named an Officer of the Order of Canada.
A founder of modern combinatorics — and a teacher for thirty years.
The cofounder
Our cofounder grew up in Cornwall, in the shadow of the Royal Albert Bridge — Isambard Kingdom Brunel’s last great work. Brunel was too ill to attend its opening; two days later, they carried him across it in an open wagon — his only crossing. He died later that year. He never got to use what he’d built — and the railway age he started rebuilt the world anyway.
Some builders never get to live in what they build. They build anyway. In 2020 our cofounder crossed an ocean in the other direction, the way Tutte did, and made Canada home. 1867 Labs is his answer to a simple question: why should the intelligence layer of a country be rented from someone else?
“Tutte was born in Britain and did his life’s work in Canada. So did our cofounder.”
Three views. Our internal synthetic suite — 7,176 items across 29 shards — the standard open benchmarks, and the attempt-by-attempt ledger: every 0.0.1 training run measured against the frozen seed, reported exactly as run. Real numbers where we have them, “pending” where we don’t — including the embarrassing ones.
MMLU, HellaSwag, ARC-Challenge, GSM8K, HumanEval — the suites everyone uses, so the numbers mean something.
| Suite | Tutte 0.0.0 (frozen) |
|---|---|
| MMLU | 24.89% |
| HellaSwag | 25.63% |
| ARC-Challenge | 25.68% |
| GSM8K | 0.08% |
| HumanEval | 1.83% |
Tutte 0.0.1 Beta scores: pending — both models run these suites via our API, so every number is comparable.
Our own training benchmark. Attempt 2’s official rework bench (Sept 24) is the run with the full per-suite breakdown — frozen 0.0.0 seed against Tutte 0.0.1 beta, reported exactly as run.
| Skill | 0.0.0 seed | 0.0.1 beta · attempt 2 |
|---|---|---|
| Overall, weighted | 29.28% | 65.23% |
| Arithmetic, two operands | 71.1% | 100% |
| Tool calling | 99.4% | 100% |
| JSON field extraction | 94.5% | 96.9% |
| Arithmetic, three operands | 0.0% | 93.2% |
| Arithmetic, multiplication | 0.0% | 98.9% |
| Copy / repetition | 0.0% | 61.5% |
| Needle retrieval | 0.0% | 0% |
| Refusal calibration — below our 60% bar | 75.0% | 8.3% |
Refusal is the known weak spot: the beta over-answers where it should decline — 8.3% in attempt 2, 16.7% in attempt 3, against our 60% bar. The ledger below tells the whole story.
Every 0.0.1 run measured against the frozen seed. Our release bar: beat 29.28% overall, move the baseline-zero suites, hold tool calling above 90%, and clear 60% refusal calibration — the gate no attempt has cleared yet. That is why nothing has shipped.
97 datasets, strictly sequential, no replay. 4.68% — catastrophic forgetting in action: JSON extraction 94.5%→0%, refusal 75%→0%. Archived as the failure that taught us the replay trainer.
The rework: family replay, retention gates, saturated waivers. 65.23% — 2.2× the baseline — with arithmetic and tool calling at 100%. No-go: refusal 8.3% against the 60% bar.
Heavier replay, ramp protocol, 97 of 97 sets. 67.95% on the full 7,176-item bench, checkpoint hash-matched to the results. Refusal climbed to 16.7% — real progress, still short of 60%. No-go.
Refusal-focused repair: probes reached 75%, but a sandbox rollback wiped the run’s artifacts. Rather than rebuild on sand, we started a fresh retrain with proper logging.
Fresh retrain from the frozen seed, full logging, official like-for-like bench running. Scorecard pending. It ships only when every gate — refusal included — passes.
We publish the 4.68% next to the 67.95% because a lab that hides its misses can’t be trusted with its hits. Attempts 1–3 are official benchmark runs recorded in the training log; the run artifacts were later lost to a sandbox rollback, so we report them as historical results rather than current releases. Attempt 4’s 75% is probe data, not a bench — marked as what it is.
The research thesis
We will never have the frontier labs’ compute, and we’ve stopped pretending otherwise. Our thesis is different: watch what the frontier discovers, adapt it fast, and automate the loop — AI agents proposing experiments, training candidates, benchmarking them, learning from the misses, with humans holding the promotion gate.
Tutte 0.0.0 was the control: an architecture, deliberately untrained. 0.0.1 is the first data point. The number we care about isn’t just the score — it’s the slope.
Small models are the point. A tiny model can be iterated thousands of times; a giant one can’t. Being compute-poor forced us to ask the more interesting question: how much intelligence can you extract from the least compute?
No alphabet soup of tiers — one model, versioned in the open. Every release stays available for comparison. Nothing is quietly retired.
The strongest Tutte yet — 67.95% on the internal suite (attempt 3) — still gated on refusal calibration. It stays beta until every gate passes. Console preview live now.
The original 1.1M-parameter mixture-of-experts seed — pinned permanently as the baseline every future Tutte is measured against. 29.28% on the internal suite. Still served for comparison.
Fresh retrain from the frozen seed — full logging, official like-for-like bench running. Scorecard pending. Ships only when every gate, refusal included, passes.
Every Tutte runs in the console with four modes: Mini for quick chats, Frontier for maximum reasoning, Uncensored for unfiltered conversation, and Era — set a year, and Tutte answers with the knowledge and voice of that time.
Tutte is small on purpose — so cheap to serve that generosity is the business model, not the marketing.
A genuinely generous free tier. A lightweight model doesn’t need a meter running — so we won’t run one.
Canadian-contracted inference for your product, priced for mortals. Build on intelligence that answers to Canadian law.
Models you can inspect, running where you say — your cloud, your country. Full Canadian hosting is the road we’re on; the mission section says exactly where we stand.
Last updated September 2026.
Frontier AI is being built by a handful of American labs. Every Canadian company that adopts it hands its intelligence layer to a foreign provider — usually without thinking twice.
Ottawa knows it: $2 billion committed (Budget 2024) to sovereign Canadian compute, with the AI Compute Access Fund now open to businesses — and AI use among Canadian businesses doubled in a year, 6.1% to 12.2% (Statistics Canada). The window where an independent Canadian lab can matter is open right now. It won’t stay open forever.
Tutte Plus and API access arrive at launch.
The console above is the preview — talk to Tutte now, no account needed. When Tutte Plus and API access open, we’ll announce it here.
We publish our benchmarks, our methods, and our misses — the ledger is the proof. And we’re building toward a published Sovereignty Standard: every Tutte release carrying a provenance statement — where it trained, where it runs, who controls it.
If you’re a Canadian university, institution, or company that wants AI you can inspect — and a lab you can actually talk to — come build with us.
This country was argued into existence in two languages. We intend to work in both — version française à venir.
Come aboard
Try the console. Read the ledger. Watch a model grow up in public.
Try TutteRead the ledger