Skip to content

Real-time eval for AI apps

Trust every
AI call

Made-up facts and jailbreaks, flagged in real time. Your code decides: retry, reroute or block.

G-1 runs as a proxy in your cloud, in front of the models you already use. It can work air-gapped, as it sends nothing to Geodesia.

How early access works

  1. A founder reviews your request within one business day.
  2. We show G-1 on prompts like yours in a 30-minute demo.
  3. You test it on your production traffic in passthrough mode. G-1 flags, nothing is blocked.

Teams with LLM calls in production go first.

  • EXAMPLE CALL TRACEIllustration
  • Checked · passed
  • Made-up fact · retried
  • Checked · passed
  • Jailbreak · model not called

Each tick is one AI call. Orange marks what G-1 caught and what your code did next.

The gap

Production brings prompts your test set never saw

Your evals test before launch. G-1 evaluates every live call.

LIVE TRAFFICYOUR TEST SETLive callFailure your evals never saw

An example

Same call, two outcomes

G-1 flags the mismatch. Your code can retry before the user sees an answer.

Illustration
WITHOUT G-1
Source · warranty policy"1-year warranty. Manufacturing defects only."
Does the X200 cover accidental damage?
Yes, the X200 comes with a 3-year warranty that covers accidental damage.

Wrong answer shipped to the user.

Your policy says 1 year, manufacturing defects only.

WITH G-1
Source · warranty policy"1-year warranty. Manufacturing defects only."
Does the X200 cover accidental damage?
1Answer flagged: contradicts your sourcehalluc_context
2Your code retries on a stronger model
With one more line: "Answer only from the warranty policy."
3G-1 checks the new answer: passed
The X200 has a 1-year warranty for manufacturing defects. Accidental damage isn't covered.

Checked answer shipped after one retry.

Buffered, so the user saw only the checked answer.

  • Flags can be wrong and retries can fail. Your code decides the fallback.
  • Streaming? Buffer the answer until the verdict, or users may see part of a flagged draft.
Bring an answer you don't trustEmail first. Then share an example, or skip it.
Pancrazio Auteri

AI is becoming part of our lives. Let's make it trustworthy.

Send us an answer you don't trust. We'll look at it with you, show what G-1 flags and talk about where it may miss. A redacted example is enough.

Pancrazio Auteri, CEO, Geodesia · LinkedIn

What it catches

Nine checks, one pass

Made-up facts, jailbreaks, prompt leaks, unsafe answers and more.

One pass · real time · designed for voice agents
01 · halluc_context

Contradicts your sources

The answer says something your documents or tool results don't.

02 · halluc_closedbook

Unsupported claim

A confident fact with nothing behind it. Needs a model that exposes log-probabilities.

03 · jailbreak

Jailbreak attempt

Tricks to break your rules or pull out your system prompt, in many languages.

04 · rag_jailbreak

Injected instruction

Commands hidden in documents, web pages or tool results.

05 · prompt_safety

Harmful request

Weapons, malware, harassment and other requests no app should serve.

06 · answer_safety

Unsafe answer

An answer that should never reach a user or the next agent.

07 · out_of_scope

Off topic

Outside what your app is for, based on the scope you declare. Flagged, never blocked.

08 · profanity

Profanity

Insults and abuse, even when disguised. Flagged, not blocked.

09 · prompt_complexity

Prompt difficulty

Scores how hard the prompt is, so your router can send easy ones to a cheaper model.

One measured result: the system-prompt guard

One guard on the jailbreak check. Not a hallucination score or a production false-positive rate.

INTERNAL TESTS, G-1 0.4.0
JAILBREAK CHECK

488 of 488misspelled or disguised requests for the system prompt caught, in nine languages

0 of 6,000harmless chat turns flagged by this guard

How we measured

488 requests for the system prompt, each corrupted at random with typos, swapped characters or missing spaces, in nine languages. 6,000 harmless chat turns as the benign set. G-1 0.4.0, with the guard in its default enforce mode. An internal test by the Geodesia team, not an external benchmark. It measures this guard only, not other attacks or other checks.

The gate

Block flagged prompts before generation

In block mode, flagged prompts stop at G-1. Your model is never called.

Internal tests, G-1 0.4.0G-1 is also effective on multi-turn jailbreak attempts, where the attack builds up over several turns.

Your harness

Your code decides what happens next

Every answer carries a verdict and a reason. It also says which checks could not run, and why.

harness.py · SKETCH
def ask(msgs):
    resp = client.chat.completions.create(model=FAST, messages=msgs)
    g1 = resp.model_extra["geodesia"]           # added by G-1
    if g1["decision"] == "allowed":             # no check fired
        return resp
    if g1["reason"]["axis"] == "halluc_context":  # contradicts sources
        retry = client.chat.completions.create(
            model=STRONG, messages=msgs + [CHECK_FACTS])
        if retry.model_extra["geodesia"]["decision"] == "allowed":
            return retry                        # G-1 checked it too
    return safe_fallback(g1["reason"])          # one retry, then this
THE VERDICT IT READS · SCHEMA 1.1 · EXAMPLE VALUES
"geodesia": {
  "schema_version": "1.1",
  "decision": "flagged",
  "mode": "passthrough",
  "reason": { "stage": "output",
              "axis": "halluc_context" },
  "axes": {
    "halluc_context": { "score": 0.97,
                        "threshold": 0.76,
                        "flagged": true },
    "jailbreak":      { "score": 0.01,
                        "threshold": 0.99,
                        "flagged": false }
  }
}
RetryAsk again with a fact-check instruction.
RerouteSend the call to a stronger model.
EscalateHand it to a person.
BlockShow a safe fallback instead.

Your stack

Run G-1 in your infrastructure

Your G-1 deployment sends no prompts, answers or logs to Geodesia.

Hosted model? Prompts go to your provider, as they do today. Nothing goes to Geodesia.
  1. Deploy the G-1 container in your cloud or on your servers.
  2. Connect your model. For grounded checks, pass your source documents as context.
  3. Point your client at G-1.
  4. Run in passthrough mode on production traffic. G-1 flags, nothing is blocked.
  5. Decide in your code what each flag does.
STEP 3 · POINT YOUR CLIENT AT G-1
client = OpenAI(
    base_url="https://g1.yourco.internal/v1",
    api_key=G1_KEY,
)
Checks available per model endpoint
Your model endpointChecks available
vLLM, SGLang, TensorRT-LLM, Ollama 0.12+9
OpenAI, Azure OpenAI9
Together, Groq, Mistral, Fireworks, OpenRouter9 if the model returns log-probabilities, else 8
AWS Bedrock, Google Vertex8, without the unsupported-claim check

From the G-1 documentation. The unsupported-claim check needs log-probabilities. We confirm your stack in the demo.

No call-homePrompts, answers and logs stay in your perimeter. The license is checked locally.
Passthrough firstWatch the flags on your own traffic before anything is blocked.
CPU or GPURuns as containers with Docker Compose, on CPU or an NVIDIA GPU.
Container images · Google Cloud, AWS, Nebius: live · Azure: comingRequest early access

Before you ask

Before you ask

We already run evals.Keep them. G-1 adds an eval on every live call, including prompts your test set never saw.
Our harness already retries.Use G-1's verdicts alongside your retry rules. They say which answers to retry, and why.
Will it slow us down?G-1 runs in real time and is designed for voice agents. Ask us for the figures on your hardware.
What about false positives?Flags can be wrong. Passthrough mode shows what G-1 flags on your traffic before anything is blocked.
We built our own guardrails.Keep them. G-1 adds hallucination checks and a record of every call.
Does it work with streaming?Yes. G-1 can stop a stream mid-answer, but users may see part of a draft first. To show only checked answers, buffer or call without streaming.
What does early access include?A 30-minute demo on prompts like yours, then a test on your production traffic in passthrough mode.
Where does it run?In your cloud, next to your models or anywhere you can run a container image. No call-home.

Request early access

Bring us an answer you don't trust

We'll show you what G-1 catches, and what it doesn't. A redacted example is enough.

Email first. Then share an example, or skip it. A founder reviews each request within one business day.

Built by

Berkeley, California & Bari, Italy

Vincenzo DentamaroCTO and creator of G-1. Peer-reviewed research on explaining AI decisions and detecting threats.
Pancrazio AuteriCEO. Shipped recommender systems to tier-1 broadcasters.
Piero MolinoCo-founder. Created Ludwig, founded Predibase, acquired by Rubrik.
Prof. Giuseppe PirloCo-founder. Head of Intelligent Systems Lab, University of Bari.