Token Trick

Human feedback · the plan

Using the data

Collect preferences from work we already needed. No duplicate runs required.

Saved today Session excerpts, your choice, your optional reason, and model and effort.

  1. Learn what matters

    Find recurring reasons behind your choices. Use them to improve agent instructions and working habits.

  2. Check whether predictions work

    Keep some comparisons out of training. Test whether an evaluator predicts your preferences on those unseen examples.

  3. Choose where to spend effort

    If predictions are reliable, use them to help choose models and effort, weighing quality alongside time and cost.

Later: Model fine-tuning, including RLHF, remains an option if the data demonstrates value.

A preference is useful evidence. Different tasks do not establish that an effort setting caused a better result, and later-discovered bugs may change the judgment.

Back to blind review