Skip to content

Evaluation

In AI analytics, I never assume an output will be reliable on the first turn, which is why I say I work slowly; I build in relevant validation of the input and verification of the output, rather than trusting either by default.

Below are four examples:

  • The vacuity index is an experiment I wrote for Research Musings, and the inspiration behind validating the input in AV|VA.
  • AV|VA is the framework the rest grew around: a proposition for proportionate validation and verification, scaled to what the LLM is being used for and what the task actually requires.
  • The verification comparison is a simplified methodology I built at work to decide which LLM we should keep internally, and to show that an AI output isn't to be trusted on the first turn, however fluent it reads.
  • Red-teaming an analytical agent is what I check when the user is adversarial rather than honest: whether a system holds its methodological and commercial integrity under pressure, not just its credentials.
Hélène reading and evaluating a document

Last updated: