Research territory

Human judgment
is measurable. Maybe.

Our research starts with a methodological problem, not a marketing claim: can longitudinal human judgment be measured defensibly across task, domain, language, rubric, context and time?

The measurement problem

Human-in-the-loop systems often treat a label, score or benchmark as if it were stable truth. We are interested in what happens when the human side is measured as a changing system instead.

That means separating accuracy, calibration, confidence, consistency, domain transfer, context sensitivity and disagreement rather than collapsing them into one reputation score.

Questions we are testing

Does judgment reliability persist across time?
How much skill transfers across domains?
When is disagreement signal rather than noise?
How do language and rubric design change measurement?
Does calibration predict useful real-world decisions?
How does AI assistance alter human judgment quality?
Can question quality be measured independently of respondent quality?
What should a multidimensional reputation system preserve?

Hypothesis ≠ finding.

We label exploratory claims as exploratory. We preserve source provenance, sample limitations, question wording and model assumptions. Synthetic experiments stay synthetic. Operational observations stay operational. Validation requires evidence that survives replication and challenge.

Prediction markets and forecasting tasks are especially useful because they create externally resolvable outcomes. They let us test whether a system that sounds intelligent is actually calibrated.

Bring us a hard measurement problem.

If your system depends on human judgment, expert review, annotation, ranking, forecasting or evaluation, the measurement layer may be more fragile than it looks.