From the open-model survey in docs/research/open-model-survey.md (item 8).
Candidate source
openjev/openjev/engine.py:84-97 _single_token_labels: probe each candidate
against a carrier string, keep it only if it adds exactly one token and that
token id has not been seen, and continue past a failure or a duplicate instead of
stopping.
What to build in riderless
validate_alphabet (riderless/api/native/worker.cpp:303-334) runs a stricter
probe than openjev's, against the real chat template and the real answer prefix,
but it breaks out of the candidate loop on the first tokenise failure and on the
first duplicate token id. On a tokenizer where one letter behaves differently
after the answer prefix, the validated alphabet silently truncates to everything
before that letter and the option capacity collapses, even though every remaining
letter is fine.
Change both break statements to skip the candidate and continue. Everything
downstream is unchanged: the alphabet is still validated at startup, the minimum
of 10 labels still applies, and the per-prompt re-check at
riderless/api/native/worker.cpp:292-299 still rejects a prompt-specific
mapping change.
Extending the candidate list past Z, which openjev also does, is explicitly NOT
part of this issue: it would change the options block for any question with more
than 26 options and is a new configuration rather than a bug fix.
Acceptance criteria
- With Gemma 4 26B-A4B the validated alphabet is unchanged: the same 26 labels in
the same order with the same token ids, so no compiled prompt moves.
- A CPU unit test covers a synthetic candidate sequence in which a middle
candidate fails to tokenise and another duplicates an earlier token id, and
asserts that the surviving alphabet is contiguous, maximal, and duplicate free.
- The startup log records any skipped candidate, so a truncation can never be
silent again.
- Guarantees to re-prove: worker code changes, so a build-to-build comparison in
the style of docs/results/worker-build-comparison.md must show every answer,
probability, and raw logit identical, and the isolation harness must still
report exactly 0.0 across its 72 comparisons.
Effort
Small.
From the open-model survey in docs/research/open-model-survey.md (item 8).
Candidate source
openjev/openjev/engine.py:84-97 _single_token_labels: probe each candidate
against a carrier string, keep it only if it adds exactly one token and that
token id has not been seen, and continue past a failure or a duplicate instead of
stopping.
What to build in riderless
validate_alphabet (riderless/api/native/worker.cpp:303-334) runs a stricter
probe than openjev's, against the real chat template and the real answer prefix,
but it breaks out of the candidate loop on the first tokenise failure and on the
first duplicate token id. On a tokenizer where one letter behaves differently
after the answer prefix, the validated alphabet silently truncates to everything
before that letter and the option capacity collapses, even though every remaining
letter is fine.
Change both break statements to skip the candidate and continue. Everything
downstream is unchanged: the alphabet is still validated at startup, the minimum
of 10 labels still applies, and the per-prompt re-check at
riderless/api/native/worker.cpp:292-299 still rejects a prompt-specific
mapping change.
Extending the candidate list past Z, which openjev also does, is explicitly NOT
part of this issue: it would change the options block for any question with more
than 26 options and is a new configuration rather than a bug fix.
Acceptance criteria
the same order with the same token ids, so no compiled prompt moves.
candidate fails to tokenise and another duplicates an earlier token id, and
asserts that the surviving alphabet is contiguous, maximal, and duplicate free.
silent again.
the style of docs/results/worker-build-comparison.md must show every answer,
probability, and raw logit identical, and the isolation harness must still
report exactly 0.0 across its 72 comparisons.
Effort
Small.