Monterey Bay AI Lab

The Evidence Agent

Making AI ocean-data answers defensible.
The trust gap in AI for science
Experts won’t recommend an AI answer they can’t cross-verify.
Temperature is the trust on-ramp — get the simplest thing right first.
The gap isn’t capability — it’s “AI-ready for what.”
Start with the simplest possible question
“What is the sea-surface temperature today?”
It looks trivial. It isn’t. A defensible answer has to resolve place, time, and what “SST” even means — and know when it can’t answer.
The agent turns one bounded question into one honest outcome
SUPPORTEDQUALIFIEDABSTAINEDSOURCE ERROR
A deterministic 15-step policy owns the disposition. No averaging into a fake “consensus.” No forecast.
1 · A defensible answer — or an honest abstention
It narrows the vague question to a bounded one, flags why the general wording is underspecified, and surfaces every comparability limit before answering.
2 · Every number traces to a raw response
Original value + units, normalized °C, request URL, retrieval time, SHA-256 of the raw response. Nothing is invented or silently substituted.
The part experts actually care about
Three feeds are not three independent votes.
The satellite analyses both blend in-situ data — so agreement isn’t independent confirmation. The tool says so, out loud.
3 · Inspect the reasoning — not a black box
Fifteen deterministic steps, each with inputs and outputs. This is an audit log, not a hidden chain-of-thought. It runs the same way every time.
4 · Correct it. Then nominate the next question.
A domain expert records a version-bound review (which never marks the result “validated”). Evidence and reviews are immutable and exportable.
Why it matters
From a confident-answer machine to a defensible-evidence machine.
The missing piece that lets experts and agencies actually adopt AI where being wrong is expensive — the answer comes with its receipts and its limits, and it knows when to stay quiet.
Honest scope
This is the trust layer, not a portal browser.
Existing portals already give breadth across thousands of layers. The hard, differentiated half — making any answer defensible — is what this prototype proves.
“The point is not that the agent always answers.
The point is that it knows what answer the evidence permits.”
Independent Monterey Bay AI Lab prototype using public data. Not endorsed by any data provider. Decision rules require domain review. Not for navigation, safety, regulatory, or operational decisions.