Appearance
Evaluation
In AI analytics, I never assume an output will be reliable on the first turn, which is why I say I work slowly; I build in relevant validation of the input and verification of the output, rather than trusting either by default.
Below are four examples:
- The vacuity index is an experiment I wrote for Research Musings, and the inspiration behind validating the input in AV|VA.
- AV|VA is the framework the rest grew around: a proposition for proportionate validation and verification, scaled to what the LLM is being used for and what the task actually requires.
- The verification comparison is a simplified methodology I built at work to decide which LLM we should keep internally, and to show that an AI output isn't to be trusted on the first turn, however fluent it reads.
- Red-teaming an analytical agent is what I check when the user is adversarial rather than honest: whether a system holds its methodological and commercial integrity under pressure, not just its credentials.
Vacuity index
An experiment I wrote for Research Musings, and the inspiration behind validating the input in AV|VA.
AV|VA
The framework those two cases sit around: a proposition for proportionate validation and verification, scaled to what the LLM is being used for and what the task actually requires.
Verification comparison
A simplified methodology I built at work to decide which LLM we should keep internally and to show that an AI output isn't to be trusted on the first turn, however fluent it reads.
Red-teaming an analytical agent
Red-teaming a pre-launch analytical agent — and the dimensions standard probes miss: whether it keeps its methodological and commercial integrity, not just its secrets.