Thede Technologies

Lab Notebook

Bake-offs, benchmarks, and build notes from Thede Technologies. The reasoning and receipts behind the tools - especially the running question of how much AI work can move on-device.

The Lab Notebook is where the methodology lives. When I run a bake-off - pitting models against a real task with a real gold set and a real judge - the scorecard goes here, sanitized and reproducible, so the conclusions are inspectable rather than asserted.

Most of these orbit one thesis I keep testing rather than claiming: how much of my daily AI work actually needs to leave my machine? Frontier cloud models still win the hard, open-ended work. But a surprising amount of the repetitive, high-volume, privacy-sensitive work - reading mail, summarizing documents, tagging a life’s worth of data - may already run well enough locally. Each entry below is one test of that question.

The ledger1 entry

No. 1

Can a Local Model Read My Inbox? A Bake-off

I gave three on-device models the same job - turn each email into a structured record - and judged them against Claude Opus. Three models, four runs, because Qwen ran once with thinking on and once with it off. With thinking off, the smallest-feeling model won on faithfulness, classification and speed.

task
Turn each email into a structured record: a summary, a type, an importance from 1 to 5, and any actions or dates
judge
Claude Opus 4.8 in the cloud, as the quality ceiling
contestants
  • Qwen3.6-35b-a3b (thinking off)
  • Qwen3.6-35b-a3b (thinking on)
  • Gemma-4-12b-it (8-bit MLX)
  • Gemma-4-e4b-it
winner
Qwen3.6-35b-a3b (thinking off)
axes
faithfulness, classification, importance, speed
verdict
Local on-device inference is good enough for this job - the winner sweeps speed and quality. The judge was a cloud model, so the test emails were read in the cloud. In daily use a model on my own machine reads the mail, and no cloud model does.

The ledger entry format

Every bake-off entry carries a structured bakeoff block in its frontmatter so the ledger index can tabulate them:

yaml
bakeoff:
  task: "<one line - what the models were asked to do>"
  date: "YYYY-MM-DD"
  judge: "<the quality ceiling / grader>"
  contestants: ["<model A>", "<model B>", ...]
  winner: "<model>"
  axes: ["<metric>", ...]      # what we scored on
  verdict: "<one-line takeaway>"

The body carries the gold set, the per-axis results, the findings, and the operational caveats that matter. Personal data used in a test never appears here - only the method and the numbers.