A small decision model for the Cardputer. It reads one typed message and answers typed questions with probabilities in a single pass, without generating any text.
The Cardputer figure scales one measured run. On a Cardputer, the ESP32 engine answered four questions about a short message, 67 tokens in all, in 2.52 s on both cores.
Options go in the second field, separated by commas; a score's levels go lowest first. Every question is answered in the same pass, and questions never see each other, so adding one doesn't change the others. To teach it, tap Mark right on the answer that should have won. Each correction stores one 128-number example in this browser. Once you have two or more, the page checks them against each other and only leans on them as much as they prove useful. Corrections shape the answers for other messages, after the page has seen a few different ones; they don't rewrite the message you corrected.
Accuracy on test sets never seen in training, scored with the same 4-bit weights this page runs. "Never trained on" rows are whole question types, held out by meaning so that no close relative was in training either. *Rows marked with a star are open, plain-language questions of the kind you would type; their answers come from a large teacher model, so they measure agreement with it, not ground truth.
| Test | Pick one | Yes / no | Score | Extract |
|---|
Across these tests, when it says 80% it is right about 80% of the time (expected calibration error, 15 bins).
On 500 pick-one questions whose right answer was removed, confidence fell from 0.58 (normal questions with three or more options) to 0.21.
12-layer encoder, 256 wide.
English only. Short typed messages work best; anything past 128 tokens is cut off. Scores and extraction are the weakest answer types. Corrections help most with categories whose names don't explain themselves; they rarely fix a wrong answer on clearly named options.