Devlog · Job Scout

Three models vote, and a register shows where they disagree

  • Triage now takes votes, not verdicts. The same batch goes to several models — here Claude, GPT and Muse — and each reply is filed under its voter's name. A voter's new reply replaces only its own earlier vote. Each unmarked job is marked from the consensus: a majority wins, and a true split (no majority) lands on Moderate, which is exactly what a split is — worth a second look.
  • Every card shows the votes as chips, one per model, with the model's reason on hover. (GIF) shows the same cards as Claude, then GPT, then Muse vote.
  • The vote register lists every voted job with each model's verdict and reason side by side, the consensus, and the person's own mark — splits first, since that's where a human call matters most. (screenshot)
  • The person still decides. A job they marked themselves keeps their mark whatever the votes say; the votes are still recorded, so the register shows where the models disagreed with them.
  • What the first real run showed. On one real day's list, three models agreed on 100 of 144 jobs, agreed two-to-one on 43, and split three ways on one. The disagreements clustered: one model consistently rated technical customer-facing architect roles higher than the other two. The register also exposed two problems in Job Scout's own batch, not the models' judgement: for some employers the trimmed description was all company boilerplate, so the models were judging titles; and some roles marked remote list only an office city. Both are now on the backlog.
Three models vote, and a register shows where they disagree, screenshotThree models vote, and a register shows where they disagree, screenshotThree models vote, and a register shows where they disagree, screenshot