Academic data is uncertain in ways that do not show up in the output.
A confidently presented figure and a heavily inferred one look identical — which makes presenting cleanly a choice, and usually the wrong one.
Where uncertainty comes from
- Identity matching — whether two records are the same person.
- Affiliation resolution — whether a name string maps to this institution.
- Coverage — what exists but is not captured.
- Timing — which date applies, and when a value was last checked.
- Source disagreement.
- None of these are visible in a clean figure.
Representing it
- Label inferred values as inferred.
- Show confidence levels where derived.
- Show last-verified dates.
- Distinguish no data recorded from no activity — these are entirely different and constantly conflated.
- Show ranges rather than point figures where the underlying data warrants it.
In aggregate figures
- State the counting rule.
- State coverage limitations alongside the number, not in a footnote.
- Say when a figure is a lower bound, which many output figures are.
- Say what would change the figure — different rule, different coverage, different date.
- An unqualified aggregate implies a precision that does not exist.
Why hiding it causes harm
- Users make decisions on figures they believe are certain.
- Sparse records get read as low activity.
- Comparisons get made between things that are not comparable.
- Errors go undetected because nothing suggests doubt.
- The harm falls on the people described, not on the data provider.
Communicating without overwhelming
- Show the essential qualifier alongside the figure.
- Put detail one click away, not three.
- Use consistent labels for confidence levels.
- Publish full methodology for those who need it.
- Uncertainty communicated well is usable; hidden it is dangerous.
One thing worth remembering
Distinguish no data recorded from no activity.
They look the same in every interface that does not deliberately separate them — and conflating them systematically misrepresents exactly the researchers and institutions whose work is least well covered.
Câu hỏi thường gặp
Where does uncertainty in academic data come from?
Identity matching, affiliation resolution, coverage gaps, timing and last-verified dates, and source disagreement — none of which are visible in a clean figure.
How should uncertainty be represented?
Label inferred values, show confidence levels and last-verified dates, distinguish no data recorded from no activity, and show ranges where the data warrants it.
What should accompany aggregate figures?
The counting rule, coverage limitations alongside the number rather than in a footnote, a statement when the figure is a lower bound, and what would change it.
Why does hiding uncertainty cause harm?
Because users decide on figures they believe are certain, sparse records get read as low activity, incomparable things get compared, and the harm falls on the people described.
Which distinction matters most?
No data recorded versus no activity — conflating them systematically misrepresents the researchers and institutions whose work is least well covered.