TUTTE 0.0.1 — API LIVE

An independent AI lab in Toronto

Built from scratch. Canadian-owned, Canadian-operated.

We’re 1867 Labs — a tiny independent lab in Toronto training its own AI models instead of renting someone else’s. Every number on this page is real, including the embarrassing ones.

We’re building toward the day every model is trained and every prompt is served on Canadian soil. We’re not there yet — and we won’t pretend otherwise. Read where we stand.

Field test

Don’t take our word for it.

The best proof is a conversation. Pick a model below and give it a hard time.

/
Preview replies · Live model API in progress
Try:··
Preview responses · Full chat at launch
The Victoria

Why build AI in Canada?

“Just use ChatGPT.”

We hear it a lot. Canada has heard it before — and answered it twice.

011864 · the crossing

Beginnings don’t get to see endings.

In September 1864, the SS Queen Victoria carried the Fathers of Confederation down the St. Lawrence to the conference that would invent Canada. The town was packed — a circus had taken every bed — so the country was argued into existence in her cabins: a floating boardroom. She sank in a hurricane off Cape Hatteras two years later — most of her crew saved by the American brig Ponvert, two hands lost — and never saw it born.

Someone still has to begin. That is what this lab is — the first crossing of Canada’s AI century.

021885 · the spine

We could have just used their railroads.

In 1871, British Columbia joined Confederation on one condition: a railway to the Pacific. Canada could have simply ridden America’s rails. Instead it laid its own — some 4,600 kilometres of steel to the Last Spike at Craigellachie in 1885 — and a scattering of provinces became a country.

Infrastructure is nationhood. Sovereignty isn’t a slogan; it’s what you build when you refuse to borrow someone else’s.

03Now · the leaf

A promise the world already trusts.

So we build our models here — owned here, answerable here. And we wrap them in the maple leaf and earn it: models people choose, a lab the world cooperates with. Not an empire play. A cooperation play.

The Victoria sails again — and this time, we finish the crossing.

The name

Named after someone who got it done.

“Some labs name their ambitions after weapons. We named ours after a codebreaker — a man who used brilliance to end a war, then spent the rest of his life teaching.”

In 1941, a 24-year-old British mathematician named W. T. Tutte walked into Bletchley Park and did the impossible: from intercepted messages alone — without ever seeing the machine — he worked out the inner structure of the Lorenz cipher, Hitler’s most secret code. It was called one of the greatest intellectual feats of the Second World War, and it helped shorten the war.

Afterward, Tutte made Canada home: professor at the University of Toronto, then the University of Waterloo, where he helped build one of the world’s great mathematics faculties. In 2001 he was named an Officer of the Order of Canada.

Tutte was British-born and Canadian-made — a deliberate mirror of this lab’s own identity. And 1867 is the year of Confederation: the bet that something built here, on purpose, can stand on its own.

1941

Broke the Lorenz cipher from messages alone. Never saw the machine.

1962

Made Canada home: professor at U of T, then Waterloo.

2001

Named an Officer of the Order of Canada.

The field

A founder of modern combinatorics — and a teacher for thirty years.

The cofounder

British-born. Canadian-made.

Our cofounder grew up in Cornwall, in the shadow of the Royal Albert Bridge — Isambard Kingdom Brunel’s last great work. Brunel was too ill to attend its opening; two days later, they carried him across it in an open wagon — his only crossing. He died later that year. He never got to use what he’d built — and the railway age he started rebuilt the world anyway.

Some builders never get to live in what they build. They build anyway. In 2020 our cofounder crossed an ocean in the other direction, the way Tutte did, and made Canada home. 1867 Labs is his answer to a simple question: why should the intelligence layer of a country be rented from someone else?

“Tutte was born in Britain and did his life’s work in Canada. So did our cofounder.”

The honest ledger.

Three views. Our internal synthetic suite — 7,176 items across 29 shards — the standard open benchmarks, and the attempt-by-attempt ledger: every 0.0.1 training run measured against the frozen seed, reported exactly as run. Real numbers where we have them, “pending” where we don’t — including the embarrassing ones.

67.95%
weighted · Tutte 0.0.1 beta · attempt 3 · best official bench
29.28%
weighted · 0.0.0 seed
7,176
benchmark items · 29 shards
Fig. 1 — standard open benchmarks

Open benchmarks

MMLU, HellaSwag, ARC-Challenge, GSM8K, HumanEval — the suites everyone uses, so the numbers mean something.

SuiteTutte 0.0.0 (frozen)
MMLU24.89%
HellaSwag25.63%
ARC-Challenge25.68%
GSM8K0.08%
HumanEval1.83%

Tutte 0.0.1 Beta scores: pending — both models run these suites via our API, so every number is comparable.

Fig. 2 — internal synthetic suite, 7,176 items

Internal synthetic suite

Our own training benchmark. Attempt 2’s official rework bench (Sept 24) is the run with the full per-suite breakdown — frozen 0.0.0 seed against Tutte 0.0.1 beta, reported exactly as run.

Skill0.0.0 seed0.0.1 beta · attempt 2
Overall, weighted29.28%65.23%
Arithmetic, two operands71.1%100%
Tool calling99.4%100%
JSON field extraction94.5%96.9%
Arithmetic, three operands0.0%93.2%
Arithmetic, multiplication0.0%98.9%
Copy / repetition0.0%61.5%
Needle retrieval0.0%0%
Refusal calibration — below our 60% bar75.0%8.3%

Refusal is the known weak spot: the beta over-answers where it should decline — 8.3% in attempt 2, 16.7% in attempt 3, against our 60% bar. The ledger below tells the whole story.

Fig. 3 — the attempt ledger

Four attempts, one bar

Every 0.0.1 run measured against the frozen seed. Our release bar: beat 29.28% overall, move the baseline-zero suites, hold tool calling above 90%, and clear 60% refusal calibration — the gate no attempt has cleared yet. That is why nothing has shipped.

Attempt 1Archived · Sept 23

97 datasets, strictly sequential, no replay. 4.68% — catastrophic forgetting in action: JSON extraction 94.5%→0%, refusal 75%→0%. Archived as the failure that taught us the replay trainer.

4.68%internal suite
Attempt 2Official bench · Sept 24

The rework: family replay, retention gates, saturated waivers. 65.23% — 2.2× the baseline — with arithmetic and tool calling at 100%. No-go: refusal 8.3% against the 60% bar.

65.23%internal suite
Attempt 3Best official · Sept 24

Heavier replay, ramp protocol, 97 of 97 sets. 67.95% on the full 7,176-item bench, checkpoint hash-matched to the results. Refusal climbed to 16.7% — real progress, still short of 60%. No-go.

67.95%internal suite
Attempt 4Superseded · Sept 25

Refusal-focused repair: probes reached 75%, but a sandbox rollback wiped the run’s artifacts. Rather than rebuild on sand, we started a fresh retrain with proper logging.

75%refusal probes
Current runIn the lab · now

Fresh retrain from the frozen seed, full logging, official like-for-like bench running. Scorecard pending. It ships only when every gate — refusal included — passes.

Pendingscorecard

We publish the 4.68% next to the 67.95% because a lab that hides its misses can’t be trusted with its hits. Attempts 1–3 are official benchmark runs recorded in the training log; the run artifacts were later lost to a sandbox rollback, so we report them as historical results rather than current releases. Attempt 4’s 75% is probe data, not a bench — marked as what it is.

The research thesis

Don’t outspend the frontier. Out-learn it.

We will never have the frontier labs’ compute, and we’ve stopped pretending otherwise. Our thesis is different: watch what the frontier discovers, adapt it fast, and automate the loop — AI agents proposing experiments, training candidates, benchmarking them, learning from the misses, with humans holding the promotion gate.

Tutte 0.0.0 was the control: an architecture, deliberately untrained. 0.0.1 is the first data point. The number we care about isn’t just the score — it’s the slope.

Small models are the point. A tiny model can be iterated thousands of times; a giant one can’t. Being compute-poor forced us to ask the more interesting question: how much intelligence can you extract from the least compute?

We optimizeΔ capability / training $·benchmark / parameter·improvement velocity

One model. An honest history.

No alphabet soup of tiers — one model, versioned in the open. Every release stays available for comparison. Nothing is quietly retired.

Tutte 0.0.1Beta · Gated

The strongest Tutte yet — 67.95% on the internal suite (attempt 3) — still gated on refusal calibration. It stays beta until every gate passes. Console preview live now.

67.95%internal suite
Tutte 0.0.0Frozen · Reference

The original 1.1M-parameter mixture-of-experts seed — pinned permanently as the baseline every future Tutte is measured against. 29.28% on the internal suite. Still served for comparison.

Next runIn the lab · now

Fresh retrain from the frozen seed — full logging, official like-for-like bench running. Scorecard pending. Ships only when every gate, refusal included, passes.

Every Tutte runs in the console with four modes: Mini for quick chats, Frontier for maximum reasoning, Uncensored for unfiltered conversation, and Era — set a year, and Tutte answers with the knowledge and voice of that time.

A model cheap enough to give away.

Tutte is small on purpose — so cheap to serve that generosity is the business model, not the marketing.

For everyone

Tutte, free.

A genuinely generous free tier. A lightweight model doesn’t need a meter running — so we won’t run one.

For builders

The 1867 API.

Canadian-contracted inference for your product, priced for mortals. Build on intelligence that answers to Canadian law.

For organizations

Sovereign deployment.

Models you can inspect, running where you say — your cloud, your country. Full Canadian hosting is the road we’re on; the mission section says exactly where we stand.

The mission

Where we stand, honestly.

True today

Canadian, end to end.

  • Canadian-owned and Canadian-operated. No foreign parent, no offshore HQ.
  • Designed with PIPEDA and Québec’s Law 25 in mind — privacy baked in, not bolted on.
  • We tell you where your data goes. No surprises in the fine print.
The road we’re on

Home soil, all of it.

  • Like most AI today, parts of our stack currently run on international infrastructure. We won’t pretend otherwise.
  • We’re moving hosting home — starting in Toronto.
  • The goal: every prompt, parameter and log on Canadian soil. Training sets rebuilt from Canadian sources.

Intelligence is concentrating.

Frontier AI is being built by a handful of American labs. Every Canadian company that adopts it hands its intelligence layer to a foreign provider — usually without thinking twice.

Ottawa knows it: $2 billion committed (Budget 2024) to sovereign Canadian compute, with the AI Compute Access Fund now open to businesses — and AI use among Canadian businesses doubled in a year, 6.1% to 12.2% (Statistics Canada). The window where an independent Canadian lab can matter is open right now. It won’t stay open forever.

Start free. Stay because it’s yours.

Free preview
$0

1,000,000 tokens. No account needed. Talk to Tutte in the console — live now.

Tutte Plus and API access arrive at launch.

Not an empire. An invitation.

We publish our benchmarks, our methods, and our misses — the ledger is the proof. And we’re building toward a published Sovereignty Standard: every Tutte release carrying a provenance statement — where it trained, where it runs, who controls it.

If you’re a Canadian university, institution, or company that wants AI you can inspect — and a lab you can actually talk to — come build with us.

Fair questions.

Why is it called Tutte?
After W. T. Tutte — the British mathematician who, at 24, cracked the Lorenz cipher at Bletchley Park without ever seeing the machine, then made Canada home: U of T, Waterloo, Officer of the Order of Canada. Some labs name their ambitions after weapons. We named ours after a codebreaker.
Aren’t your benchmark scores… kind of bad?
Compared to the giants? Yes. We’re a tiny lab training from scratch, and we’re showing you the real numbers anyway — because a lab that only publishes its wins is a marketing department. Watch the ledger move.
What is 1867 Labs?
1867 Labs is the lab; Tutte is the model. 1867 Labs is Canada’s sovereign AI lab — designed, owned and operated in Canada. Tutte is our AI, versioned in the open. And 1867 is Confederation: the bet that something built here, on purpose, can stand on its own.
Does my data stay in Canada?
Not entirely — not yet. Like most AI today, parts of our stack currently run on international infrastructure, and we won’t pretend otherwise. Full Canadian residency is the road we’re on: hosting moving home starting in Toronto, training sets rebuilt from Canadian sources. This page says where we stand, and we update it as that changes.
How is this different from the big US labs?
Ownership and accountability. 1867 Labs answers to Canadian law, a Canadian team, and Canadian users — not a foreign board. Sovereignty is who answers for your data.
Is my data used to train the models?
Our policies are plain-language and transparent — no surprises buried in legalese. Full details ship with the launch.
When is the API coming?
It’s on the roadmap — we’ll announce it here the moment API access opens.
« Construire ici. »

This country was argued into existence in two languages. We intend to work in both — version française à venir.

Come aboard

The Victoria sails again.

Try the console. Read the ledger. Watch a model grow up in public.

Try TutteRead the ledger