jev for seo is insane (here's why)

Jev launched on September 15 behind a waitlist. It doesn't chat. It doesn't generate text. You send it state and a typed question, it returns a probability in about 250 milliseconds, for a fraction of a cent.
I think it changes how SEO gets done at scale. Here's why, with every request written out so you can run it the day you get in.
Full disclosure first: I don't have access yet.
I'm on the waitlist like everyone else.
Everything below is researched.
Oh and before we start, if you're struggling with SEO/GEO
--> This agent will make your life easier (and get you customers)
— "chatseo.app"
TLDR: An LLM talks. Jev decides.
Jev is the first model from TypeSafe, a San Francisco lab founded by Diogo Almeida, who worked at OpenAI on the instruction-following methods behind ChatGPT. They call it a System One model, after Kahneman's fast, intuitive thinking. No chat interface, no chain of thought. It reads your input and puts probability mass on the answers you allowed.
Three consequences for SEO:
- It cannot answer outside the schema. Ask "commercial, informational, transactional or navigational?" and you get one of the four, with a probability each. No fifth answer, no preamble, no JSON that fails to parse on row 4,812.
- It's fast because it doesn't write. TypeSafe quotes 70 to 500ms end to end. Third parties measured a median around 244ms. Questions in one request run in parallel, so asking forty costs about the same latency as asking one.
- It's priced like a database query. $0.042 per million input tokens, output free. Judging a 50,000-keyword export costs a dime. Sampling 500 and guessing the rest stops being a rational trade-off.
Caveat that preempts your reply: the speed and cost comparisons are TypeSafe's own numbers, vendor-run and unreproduced. The pricing and latency ranges are published and third parties confirm the latency. Treat the "400x cheaper than LLMs" framing as marketing until independent benchmarks exist.
The three question types
Every request is one state plus named questions. Three kinds, and every SEO judgment below is one of them:
Type Asks Returns SEO use noul Is this statement true? Probability of yes, 0 to 1 Keep or prune. Same intent or not. choice Which of these options? Choice + probability per option + confidence Intent class. Funnel stage. Which page wins. score Where on this rubric? Score + probabilities + confidence Draft quality. Cannibalisation severity.
The names are theirs. Two fields do the real work: the probability tells you the answer, the confidence tells you whether to act without a human. TypeSafe's guidance: act automatically above 0.9, cautiously between 0.5 and 0.9, route to a person below 0.5.
That's the whole architecture, and it's the opposite of how most people use AI for SEO. You don't ask Jev to "analyse my keywords". You keep the loop in a script, hand Jev one narrow question per row, and let thresholds in your own code decide what happens next. Every SEO tool built on chat prompts got this backwards.
What it's bad at, from TypeSafe's own docs
The docs ship a page called "jaggedness" and it changes how you write questions:
- It answers the question you wrote, not the one you meant. "Is this query NOT relevant" will burn you. Phrase every noul so a high value means yes to something positive.
- It cannot count. Word counts, H2s, occurrences: code counts, Jev judges.
- It reads numbers as text. "Did position improve" is a spreadsheet question. Don't ask it.
- Irrelevant state is a distractor. Accuracy falls as the state fills with things unrelated to the decision.
- "Cannot hallucinate" means it can't answer outside your schema. It can still pick the wrong option inside it. The confidence field tells you when.
The four workflows
Each one: a state, a question, a threshold, one action.
- Filter your full GSC export down to relevant queries
A 16-month Search Console export is mostly noise. Brand misspellings, accidental rankings, 2,000 long-tails with one impression. Every keyword analysis you've ever done was done on a sample because the full list was unreadable. Jev reads the full list.
State: two or three sentences describing the business, plus a batch of 200 queries keyed by id. One noul per query:
{
"model": "jev-latest",
"state": {
"site": "ChatSEO is an AI SEO assistant for small business owners and solo founders. It connects to Google Search Console and tells the user which page to fix next.",
"queries": {
"q1": "chatgpt prompts for seo",
"q2": "netflix cancel subscription"
}
},
"questions": {
"q1": {
"type": "noul",
"instructions": "Someone typing the query at queries.q1 into Google is a plausible customer for the business described in site.",
"criteria": {
"true": "The query is about doing SEO, understanding Search Console data, or choosing an SEO tool, from someone who runs a website.",
"false": "Unrelated topic, hiring an agency, academic, or navigational for another product."
}
}
}
}
Thresholds: ≥ 0.8 keep, ≤ 0.2 drop, in between tag for review and read the 40 highest-impression ones by hand.
One action: the keep list, joined back to clicks and impressions, becomes your working keyword set. Workflows 2 to 4 run on that set, never on the raw export.
- Classify intent and funnel stage into YOUR taxonomy
Every intent classifier you've used either forced Google's four buckets on you or let an LLM invent a sixth category on row 300. A choice question is closed. Your options, a probability per option, no escape.
Two choice questions per query: one for intent (informational, commercial, transactional, navigational, other), one for funnel stage (top, middle, bottom, none). Always include an "other" option, or ambiguous queries get forced into your best-sounding bucket with fake certainty.
One action: pivot by stage × intent, filter to bottom + transactional, sort by impressions, find the first query whose ranking URL is a blog post or empty. That's the page you brief this week. Not the list. That one.
- Detect cannibalisation across every page pair
The usual check is "which queries have two URLs". Half of those are fine. The real question, same intent or not, is a judgment that was too expensive to make for every pair. So nobody did. At this price you can.
Code pre-filters first: from the query + page export, build pairs that share at least one query with impressions. A few hundred pairs, not 79,800. Then three questions per pair: same_intent (noul, the decision), winner (choice: a, b, or neither, the plan), severity (score, the priority). State is title, H1, meta and top five shared queries. Not the full page bodies. That's exactly the irrelevant state the docs warn about.
Merge candidate when same_intent ≥ 0.9, severity ≥ 1.5, winner confidence ≥ 0.7.
One action: the top pair gets consolidated. Loser 301s to winner. Then you measure for 28 days before touching the next one, because "we merged 15 pages in a weekend" is how people find out their cannibalisation checker was wrong.
- Score AI drafts before publishing
If you generate content at volume, the bottleneck moved from writing to reviewing a long time ago. A frontier model as reviewer is slow, costs more than the draft did, and gives a different verdict Monday than Tuesday. A reviewer that changes its mind between runs is not a gate, it's a coin.
Per draft: matches_intent (noul), coverage of the brief (score), unsupported_claim (noul), sounds_generated (score), verdict (choice: publish, revise, rewrite).
Auto-approve only when verdict = publish with confidence ≥ 0.9, matches_intent ≥ 0.9 AND unsupported_claim ≤ 0.1. Everything else goes to a review queue with the failing question's name attached.
Note what none of these ask: word count, H2 frequency, keyword in first 100 words. Jev can't count. Those are five lines of code that run before the request.
The test that matters
Whatever backend you use: hand-label 50 rows first. Run the workflow. Count the disagreements. Under 5, trust the ≥ 0.9 band. Over 10, rewrite the question, not the threshold.
That 50-row check is the difference between using Jev and trusting a number you never verified. A calibrated 0.92 is only calibrated on the distribution it was trained on. Your queries aren't that.
You can build this today, without access
The API is waitlisted, but the wire format is public and already cloned. Three routes:
- TypeSafe direct. Waitlist at console.typesafe.ai.
- Vercel's AI Gateway. Lists Jev as a model, billed through your Vercel account. The route the first third-party libraries default to.
- Open-weight clones on your Mac, today. Kev (Qwen3, 32GB Mac), Laya (ModernBERT, ~10ms on Apple Silicon), open-jev (Gemma 3 4B). Not Jev, not as smart, same request shape. Write and debug every workflow now, swap the base URL when your key arrives.
The mistakes I expect to make first
Written before I've made them:
- Trusting the probability without the 50-row check.
- Asking a negated question. "Is this irrelevant" fails in a way you won't notice for a month.
- Stuffing state. Title, H1, meta, top queries. Stop there.
- Asking Jev to count or compare numbers.
- Acting on the whole list instead of one action per workflow. It's the only way to learn whether the judgment was right before you've made it four hundred times.
Where this stops
Every request above judges. None of them sees. Jev doesn't pull your Search Console, doesn't know which queries gained position since last month, doesn't check 28 days later whether the merge worked. That loop is still you, a script and a spreadsheet.
That's the gap ChatSEO closes. It's connected to your Search Console, so the judgment lands on live numbers, the action gets pushed to your CMS, and the recheck is on the calendar. The day Jev is in its toolbox, these four workflows are what it runs underneath. Until then, they're what I'll run by hand.
The judgment just got cheap. The action didn't. The action is still the part that moves clicks.
Related articles

How We Generated Millions Of Organic Clicks With Programmatic SEO
most SEO looks like this:

Our Entire $35M AI SEO Playbook (and how to copy it)
in 2003 you could publish a blog post and rank on Google in a WEEK

I told ChatGPT-6 Astra to do my SEO. It got me $27k MRR (full guide)
Connect a custom MCP server named Conqueror using this URL: [https://winwith.conquerorapp.com/mcp](https://winwith.conquerorapp.com/mcp)