Real-time eval for AI apps
Trust every
AI call
Made-up facts and jailbreaks, flagged in real time. Your code decides: retry, reroute or block.
G-1 runs as a proxy in your cloud, in front of the models you already use. It can work air-gapped, as it sends nothing to Geodesia.
How early access works
- A founder reviews your request within one business day.
- We show G-1 on prompts like yours in a 30-minute demo.
- You test it on your production traffic in passthrough mode. G-1 flags, nothing is blocked.
Teams with LLM calls in production go first.
- EXAMPLE CALL TRACEIllustration
- Checked · passed
- Made-up fact · retried
- Checked · passed
- Jailbreak · model not called
Each tick is one AI call. Orange marks what G-1 caught and what your code did next.
The gap
Production brings prompts your test set never saw
Your evals test before launch. G-1 evaluates every live call.
An example
Same call, two outcomes
G-1 flags the mismatch. Your code can retry before the user sees an answer.
Wrong answer shipped to the user.
Your policy says 1 year, manufacturing defects only.
Checked answer shipped after one retry.
Buffered, so the user saw only the checked answer.
- Flags can be wrong and retries can fail. Your code decides the fallback.
- Streaming? Buffer the answer until the verdict, or users may see part of a flagged draft.

AI is becoming part of our lives. Let's make it trustworthy.
Send us an answer you don't trust. We'll look at it with you, show what G-1 flags and talk about where it may miss. A redacted example is enough.
Pancrazio Auteri, CEO, Geodesia · LinkedIn
What it catches
Nine checks, one pass
Made-up facts, jailbreaks, prompt leaks, unsafe answers and more.
Contradicts your sources
The answer says something your documents or tool results don't.
Unsupported claim
A confident fact with nothing behind it. Needs a model that exposes log-probabilities.
Jailbreak attempt
Tricks to break your rules or pull out your system prompt, in many languages.
Injected instruction
Commands hidden in documents, web pages or tool results.
Harmful request
Weapons, malware, harassment and other requests no app should serve.
Unsafe answer
An answer that should never reach a user or the next agent.
Off topic
Outside what your app is for, based on the scope you declare. Flagged, never blocked.
Profanity
Insults and abuse, even when disguised. Flagged, not blocked.
Prompt difficulty
Scores how hard the prompt is, so your router can send easy ones to a cheaper model.
One measured result: the system-prompt guard
One guard on the jailbreak check. Not a hallucination score or a production false-positive rate.
INTERNAL TESTS, G-1 0.4.0
JAILBREAK CHECK
488 of 488misspelled or disguised requests for the system prompt caught, in nine languages
0 of 6,000harmless chat turns flagged by this guard
How we measured
488 requests for the system prompt, each corrupted at random with typos, swapped characters or missing spaces, in nine languages. 6,000 harmless chat turns as the benign set. G-1 0.4.0, with the guard in its default enforce mode. An internal test by the Geodesia team, not an external benchmark. It measures this guard only, not other attacks or other checks.
The gate
Block flagged prompts before generation
In block mode, flagged prompts stop at G-1. Your model is never called.
Internal tests, G-1 0.4.0G-1 is also effective on multi-turn jailbreak attempts, where the attack builds up over several turns.
Your harness
Your code decides what happens next
Every answer carries a verdict and a reason. It also says which checks could not run, and why.
def ask(msgs):
resp = client.chat.completions.create(model=FAST, messages=msgs)
g1 = resp.model_extra["geodesia"] # added by G-1
if g1["decision"] == "allowed": # no check fired
return resp
if g1["reason"]["axis"] == "halluc_context": # contradicts sources
retry = client.chat.completions.create(
model=STRONG, messages=msgs + [CHECK_FACTS])
if retry.model_extra["geodesia"]["decision"] == "allowed":
return retry # G-1 checked it too
return safe_fallback(g1["reason"]) # one retry, then this"geodesia": {
"schema_version": "1.1",
"decision": "flagged",
"mode": "passthrough",
"reason": { "stage": "output",
"axis": "halluc_context" },
"axes": {
"halluc_context": { "score": 0.97,
"threshold": 0.76,
"flagged": true },
"jailbreak": { "score": 0.01,
"threshold": 0.99,
"flagged": false }
}
}Your stack
Run G-1 in your infrastructure
Your G-1 deployment sends no prompts, answers or logs to Geodesia.
- Deploy the G-1 container in your cloud or on your servers.
- Connect your model. For grounded checks, pass your source documents as context.
- Point your client at G-1.
- Run in passthrough mode on production traffic. G-1 flags, nothing is blocked.
- Decide in your code what each flag does.
client = OpenAI(
base_url="https://g1.yourco.internal/v1",
api_key=G1_KEY,
)| Your model endpoint | Checks available |
|---|---|
| vLLM, SGLang, TensorRT-LLM, Ollama 0.12+ | 9 |
| OpenAI, Azure OpenAI | 9 |
| Together, Groq, Mistral, Fireworks, OpenRouter | 9 if the model returns log-probabilities, else 8 |
| AWS Bedrock, Google Vertex | 8, without the unsupported-claim check |
From the G-1 documentation. The unsupported-claim check needs log-probabilities. We confirm your stack in the demo.
Before you ask
Before you ask
Request early access
Bring us an answer you don't trust
We'll show you what G-1 catches, and what it doesn't. A redacted example is enough.
Email first. Then share an example, or skip it. A founder reviews each request within one business day.
Built by
Berkeley, California & Bari, Italy



