Product case study · Jacob Dixon
A five-handed Sheepshead game for a group of friends who could never all be free on the same Thursday — and a working demonstration of how I run product: research before sequence, measurement before belief, and every reason written down where the next person can check it.
Live at noschnitz.com. Solo against four AI opponents, or multiplayer with a shareable link. Figures below are as of v0.120.1.
01 — Where the roadmap came from
The group had already tried this. Understanding why it failed was the whole roadmap.
During the pandemic the same five-to-seven people ran a biweekly Thursday Sheepshead night on an existing site, with a separate video call alongside it, coordinated over group text. It failed more often than it worked.
The interview found that it wasn't failing for lack of interest. It was failing because a scheduled commitment kept losing to spontaneous life, and getting five people to commit in advance was the actual bottleneck. On the good nights the opposite problem appeared: six or seven people wanted in, and the tool couldn't gracefully rotate anyone through five seats.
The finding that reordered everything was smaller and less flattering to the product: the video call was doing the emotional work. The card game was an excuse to fill the quiet spaces.
That became a written hypothesis at the top of the roadmap, with everything in Now scoped to test it and nothing else: will this group actually play together, on a whim, without a scheduled event? AI-filled seats so there is no minimum headcount. A shareable link so nobody has to cajole a group text. Presence so "is it worth hopping on" answers itself.
A roadmap that starts from a hypothesis can be wrong in public. Six days in, the first real session ran two humans and three AI seats through roughly twenty hands with no stalls — the table filled to five with only two people in the room, which is precisely the bet. The write-up in the repo says so and then immediately says not to read it as coverage: two humans is not five, nobody joined mid-hand, nobody backgrounded a phone. A result you are willing to under-claim is a result you can still learn from.
02 — Sequence
A board shows what is being worked on and never why that rather than something else.
Features decompose into epics and epics into user stories, with stable IDs so a commit can name the story it implements. But the part that does the work is the paragraph under each bucket explaining why the order is the order — because a reader who can't see the reasoning concludes the sequencing was accidental rather than chosen.
Scoped to test one hypothesis: will they play on a whim, unscheduled?
Earns its place only once the Now hypothesis has an answer.
Deliberately deferred — with the reason attached, so deferral stays a decision.
The tutorial is core to the long-term goal and sits in Later anyway, because it can't pay off until public tables exist. Deferring it would normally mean quietly killing it — so the deferral came with an architecture constraint written into the same document: multiplayer had to be built as an additional mode alongside solo, reusing the pure rules engine, rather than a rewrite that entangles local and networked state. The tutorial slots in later exactly the way solo works today.
Sequencing a thing later is easy. Keeping it cheap to build later is the actual work.
03 — A decision, in full
The number was unambiguous. It was also answering a different question than the one that mattered.
In Sheepshead the picker can go alone — no partner, quadruple stakes. The engine holds a threshold for when the AI takes that option. The app holds a second threshold for when it offers the option to a human. There is no rule that says they must be the same number.
Across 20,239 hands where the picker had a partner available, going alone at a threshold of 17 was behind by 1.9 points per hand, negative in four of four seeds. It only turned positive at 18. The offer bar was set to 18 on the strength of exactly that.
Then somebody played the game.
Shipped: 17.v0.46.0
Overriding a measurement is only defensible when you say what it costs, write the condition that would reverse you, and make the codebase fail loudly if someone later moves half of it.
04 — Evidence
Every one of these rules exists because a specific measurement was wrong first.
Engine and AI changes are decided by paired A/B runs: a variant in one seat against four unchanged seats over identical shuffles, across several seeds. What makes a small result trustworthy is the null control — an arm identical to the baseline must return exactly +0.0000, and the test suite asserts it. A harness that can't return zero when nothing changed can't be believed when something does.
That was the easy part. The hard-won part is the standing list of ways the instruments mislead:
Aggregate simulation is the safety net, not the detector. Every AI fix that has actually landed started from one hand a human flagged, was reproduced against the engine before anything changed, and was then pinned as an assertion with a negative control. Several correct-looking diagnoses measured as pure noise.
05 — Cadence
Real data from the changelog. Each bar is one day; each unit is one versioned release with a written rationale.
06 — The operating system
Two AI agents work this repo under a written boundary — one owns the roadmap and the board, one owns the code, and I own the decisions. Everything below exists because that only works if the contract is explicit.
The backlog started as a markdown file. Within twenty-four hours it had collected a duplicate heading, two shipped items still open, and a line describing a deleted file as present in the working tree. It was retired for the issue board — not on taste, but on the observation that every mechanism a shared mutable file needs — section ownership, a claim protocol, re-read-before-closing, batching to control cost — is scaffolding for a merge conflict. Issues have no merge.
07 — The uncomfortable one
The most useful document in this project is the one that audits the rest of it.
Partway through, the board got read the way an outsider would read it. The finding was structural rather than careless, and it is the one I'd want a hiring committee to see.
The AI pick threshold — an item worth a fraction of a point per hand — had a null control, sign agreement across splits, a throw-in budget check and a control arm under a second rule set. Real-device multiplayer coverage sat in Later with no acceptance criteria at all and a note that it was "largely covered in practice." By the project's own written evidence, every multiplayer bug worth fixing had come from a person on a phone, and the newest, least-exercised surfaces had never been touched by five people at once.
The output wasn't a reshuffle. It was a standing question, written into the conventions where it gets asked again: is this rigorous because it matters most, or because it was the easiest thing to be rigorous about?
The related test for whether reasoning is genuinely written down: hand the board to someone who wasn't in the conversation. Everything they need is reachable from a clone. Every number has one home. Every gate can be evaluated by someone who wasn't there. Every issue says what not to do. Anything unshaped is labelled unshaped. Where that fails, the reasoning is still in somebody's head — with a paper trail that looks like documentation, which is the more expensive failure, because it's the one nobody checks.
08 — Constraints
The API sits at eleven of the hosting plan's twelve serverless function slots, and the limit only fails at deploy time. That single fact is why "collapse the action endpoints" is a filed, prioritised backlog item framed as what makes room to build anything rather than as cleanup — and why any change that wants a new route reads that constraint first.
Deployments, not CI minutes, are the scarce resource: the plan allows a hundred a day, and one day of two-agent work spent well over thirty. So the deploy cost of every branch is written down as a table — zero per push to a working branch, one per merge to the integration branch, one per promotion — and documentation-only changes skip the build via a version-controlled ignore rule whose every failure direction is toward building rather than toward a missed deploy.
For most of the project every merge shipped to production. Once real people were playing between releases, that stopped being right: an integration branch went in the middle, and promotion became a deliberate, milestone-driven act with an explicit merge strategy — because squashing a promotion loses the merge base and makes the next promotion re-conflict everything already shipped. The old cadence isn't deleted from the record; it's marked historical, so a future reader knows which instructions are current.
09 — What this is evidence for