how to run an ai native agency with jev (FULL GUIDE)

jev is open to everyone now. it's a super fast decision model, so one person can now do work that businesses usually pay an agency a big cut for. here's how you can set up a whole ai native agency with it.
by the end of this guide you'll have:
- what jev is and how to get access to it
- how an ai native agency works and gets paid
- three services businesses already pay a big cut for
- the exact rules and prompts to run each one
- real accuracy numbers from my own run of all three
- a demo page you can send to any business
- all of it running on hermes or grok bot

what jev is
when you ask chatgpt or claude something, you get sentences back. jev gives you a decision instead. you give it a situation and a short list of options, it picks one, and it tells you how sure it is.
let's say you show it an unpaid invoice and ask whether to send a reminder, send a firmer message, or stop because the customer already paid. it picks one and tells you how sure it is.
when jev is sure, the work goes out on its own. when it isn't, it comes to you. so basically, you're not checking everything it does, just the few cases it isn't confident about.
and because it can only pick from the options you gave it, it can't make something up. it can't invent a payment that never happened or a fee that isn't on the bill. it just picks.
it answers in under a second and costs a fraction of a cent per decision, so running thousands of checks a day costs very little.
to get access, sign up at console.typesafe.ai and create an api key under api keys, or use openrouter with the model name ~typesafe/jev-latest if you already have an account there. an api key is just a password that lets your code use the model.
what an ai native agency is
every small business is basically owed money it never gets back. customers who didn't pay, bank disputes they lost, bills they got overcharged on. there are agencies that recover that money and take a cut.
most small businesses already pay for this. they just pay a lot for it, and the agencies doing it are slow, because a person has to go through every invoice, every dispute and every bill by hand.
that work is mostly checking things against a set of rules, which is exactly what jev is good at. so one person can basically be that agency for fifty businesses.
the difference between this and a normal agency is what you sell. a normal agency sells you their time. an ai native agency sells the finished result, so the business doesn't pay for hours, they pay when the money actually comes back.
you write down the rules for what the right call is, jev checks every case against them, and you get paid a cut of whatever money comes back.
the five main pieces
every service in this guide is built from the same five parts.
the unit is one thing you work on, so one invoice, one dispute or one bill. it's what you price by, and it's what the business thinks in, so keep it simple.
the intake is the business sending you the pile, usually an export from their accounting software, their stripe or shopify dashboard, or a folder of pdf bills. it doesn't need to be clean and structured, the agent can read whatever they send.
the engine is fable, which writes whatever goes out, like the reminder email or the letter. jev decides what should happen and fable writes it, so each one is only doing the thing it's good at.
the rulebook is your written list of what the right call is in each situation. this is the part you build up over time, and it's the part that makes your agency better than the next one.
the review layer is where anything jev isn't sure about comes to you for approval. this is what lets one person run it for lots of businesses, because you only ever look at the few it's unsure on.
writing the rulebook
every service starts with you writing out in plain words what the right call is in each situation. a list of sentences is fine, you don't need any special format.
the easiest way to write one is to imagine you're training a new hire on their first day. what would you tell them to do in each situation, and when would you tell them to stop and ask you? write that down.
a good rule is specific enough that two people reading it would make the same call. "chase late invoices" is too loose. "send reminder 1 at 7 days overdue" is a rule. the more of your rules look like the second one, the fewer cases end up coming to you.
then claude turns your rules into the questions jev asks about every case.
turn this rulebook into jev questions.
- here's the rulebook in plain words: [paste your rules]
- for each decision the rules describe, write one jev question with its list of possible answers
- use pick one from a list when there are several possible actions, a scale when there's an order like low, medium, high, and true or false for yes or no checks
- write the confidence line for each one, what runs over 0.8, what goes to review below it
- anything that must always stop, like already paid or past a deadline, is its own true or false question that overrides everything else
- tell me which of my rules are unclear enough that jev would have to guess, and ask me about them before you write the questions

service one, overdue invoices
every business has customers who pay late, and today they pay a collections agency somewhere between 15 and 30% of whatever gets recovered.
you can charge less, say 5% of what comes in, and you start chasing at 7 days, when the customer is still likely to pay. (you can charge however much you'd like)
the business sends you their unpaid invoices and each customer's history, like any replies or promises to pay. that history stops jev chasing someone who already paid, or someone who's already said they'll pay on friday.
the earlier you catch them, the more of the money actually comes back, and the friendlier the message can be.
here's the rulebook, written the way you'd explain it to a new hire.
send reminder 1 at 7 days overdue
send reminder 2 at 21 days
send a firm notice at 45 days
call at 60 days
stop if the customer disputed the invoice
stop if they promised a payment date that hasn't passed yet
stop if they're on a payment plan
stop if the amount is under $50
never chase an invoice that's already paid
never chase the same customer twice in 7 days
jev decides which step each invoice is on and whether it should stop, and picks a tone. so a good customer who's 8 days late gets something friendly, and someone who's ignored two reminders gets something firmer. fable writes the message.
run the invoice rulebook against this pile.
- here are the invoices: [paste or attach]
- for each one, send jev the invoice, the customer history and the rulebook questions
- apply the stop rules first. any stop over 0.8 wins
- if the action is over 0.8, have fable write the message in the tone jev picked
- everything under 0.8 goes into a review list with the reason
- show me a table of every invoice, the action, the probability and whether it went out or to review
- don't send anything. prepare the batch and stop
what happened when i ran it
i ran it on 60 invoices for a plumbing company in phoenix. it got 96% right, gave the same answers both times i ran it, and never once chased someone who'd already paid.
getting chased for something you've already paid is the mistake that actually annoys a customer, and that never happened. the handful that came to me for review were customers with a messy history, like a part payment and a reply in the same week, which is exactly what i'd want to look at myself.
the one mistake it was confident about was a customer who'd replied "what's this one for again?". jev read that as a dispute and stopped chasing. but they weren't disputing anything, they were asking a question in this situation.
so i added one line to the rulebook: if the customer asked a question, answer it first. that's basically how the rulebook grows, one mistake at a time.

service two, chargebacks
a chargeback is when a customer asks their bank to reverse a payment instead of asking the business for a refund. shops, restaurants and online stores lose a lot of these, mostly because nobody answers them properly before the deadline. agencies fight them for 20 to 30% of whatever's won.
you charge a cut of what you win back. once the business says yes, they connect their stripe or shopify so you can see the disputes.
each dispute comes with a reason, like the customer saying they never got the item, or that they didn't recognise the charge, or that the product wasn't what was described. what wins depends on the reason, so the rulebook is organised by it.
every dispute also has a deadline. miss it and the business loses automatically, whatever proof they had, so that always comes first.
for each reason code, the evidence that wins:
delivery confirmation to the billing address
a signed receipt
address and cvv match on the card
the customer's own messages
the refund policy shown at checkout
proof the customer used the product or service
fight if the winning evidence is present
let it go if it isn't
always respond before the card network's deadline
jev decides whether to fight, let it go, or gather more evidence, and it only fights when it's sure there's a real chance of winning. when it fights, fable writes the dispute packet, which is the document the business sends back to the bank with the proof attached.
run the chargeback rulebook against these disputes.
- here are the disputes with their reason codes and the evidence we have: [paste or connect]
- for each one, send jev the dispute, the reason code, the evidence and the rulebook questions
- if it's a fight over 0.8 with at least medium likelihood, have fable write the dispute packet
- if the deadline is at risk, put it at the top of the review list whatever else jev said
- everything else goes to review with the reason
- show me every dispute, the decision, the probability and the deadline
- don't submit anything. prepare the packets and stop
what happened when i ran it
30 disputes for a home goods store in phoenix. it got 90% right, and it never fought one it couldn't win or gave up on one it could.
the few it got wrong, it asked for more evidence on, which just means they came to me. that costs nothing.
the worst mistake here would be fighting a dispute you can't win, or giving up on one you could have won, and it did neither. everything that went through came with a finished packet ready to send.

service three, bill audits
businesses overpay on card processing, phone and internet, and waste collection, mostly because nobody checks the bill against what they agreed to pay. audit companies find these and take 30 to 50% of the savings.
you take a share of what you save them, so it's free until it works.
the business sends you their bills and their contracts. the contract is just the document that says what they agreed to pay. a lot of businesses can't find theirs straight away, and that's fine, the agent compares against the provider's public prices and their older bills instead until it turns up.
most of these overcharges are small on their own. a fee that crept in, a rate that went up without anyone noticing, a service they cancelled but are still paying for. but they're on every bill, every month, so they add up to real money over a year.
flag a line as an overcharge if it is:
a fee that isn't in the contract
a rate above the contracted rate
a charge for a service they no longer use
the same fee charged twice
a rate increase with no notice given
a minimum charge applied when their volume was over the minimum
compare each line against the contract, the published rate table,
and the last three bills
jev goes through every line on the bill and flags the ones that are wrong. then the total gets added up and fable writes the letter asking for the money back.
the adding up is done separately from jev on purpose. jev decides which lines are wrong, and simple maths works out the total, so the number is exactly the same every time you run it. that's the number you show the business, so it has to be right.
run the bill audit against these bills.
- here are the bills and the contract: [attach]
- for each bill, send jev every line, the contract terms, the published rates and the last three bills
- flag a line when jev is over 0.8 sure it's an overcharge, and record which type
- add up the flagged lines in code into a monthly total and an annual total
- have fable write one dispute letter per bill with every flagged line listed
- anything jev named correctly but under 0.8 goes to review
- show me every line, the flag, the type and the probability
what happened when i ran it
20 bills for an hvac company in phoenix. it caught 34 of the 35 overcharges i'd hidden in them, never flagged a line that was actually correct, and found $22,047.96 a year they were owed.
it also left alone a couple of charges that looked wrong but weren't, like a fuel surcharge that matched the provider's published table. that's just as important as catching the real ones, because one wrong flag in a letter makes the business look bad.
the one it missed was a duplicate charge. it spotted it and named it correctly, but it wasn't sure enough, so it came to me instead of going in the letter.

the rulebook fixes itself
across the three runs, jev and i disagreed twelve times, and more often than not jev was the one who was right.
the invoice question was jev being wrong. but on the chargebacks, i'd been following a rule in my head that i never wrote down, so jev couldn't have known it. the fix was writing it down.
another time, my rulebook asked for a letter as evidence that the business only writes when they submit, so it couldn't exist yet. jev kept saying the evidence was incomplete, and it was right.
whenever you disagree, look at the case, decide who was right, and fix either the rule or your answer. then run it again and check it got better.
after a few rounds of this, the disagreements get rarer, because every one of them has turned into a written rule. your rulebook ends up covering the odd cases that only come up when you run it on real work.
every one of those became a line in the rulebook. that's how you keep improving the system.

how you sell it
you build a demo page for the business before you contact them.
it has their real name and logo at the top, then made up example data in their name, like ten invoices, three disputes or one bill. next to each one is what jev decided, and under the top ones is the actual message or letter it wrote. at the bottom is how much money it found, clearly marked as example data.
then one line: send us your 20 oldest invoices, your last 3 disputes, or one bill, and we'll run it for free.
the free first job gets you their real data. when they send it, you run it exactly the same way and send back the results, with anything jev wasn't sure about already checked by you.
for invoices, that's a batch of messages ready to go. for disputes, it's the packets. for bills, it's the letter and the total. once they see a real number from their own business, you agree the cut and turn it on for everything going forward.
build a demo page for [business] for the [service] service.
- their real business name and logo at the top
- mock data in their name: 10 invoices, 3 disputes, or 1 bill
- run it through the rulebook and show jev's decision and probability next to each item
- show the message, packet or letter for the top items
- the annual total at the bottom, labelled clearly as mock data
- one line at the end offering to run their real data free, with a link to send it
- keep it to one page that works on a phone
- don't use any real figures from their business. if you can't tell whether a number is real, leave it out
the page goes on a postcard with a qr code, or straight into a dm with the link.

what it costs to run
my whole test across all three services cost about 7 cents. so running this for a customer costs you less than a dollar a month, and you're charging a cut of thousands.
that also means you can run the free first job for anyone without thinking about it. it costs you almost nothing to show a business what they're owed.
running it on hermes or grok bot
once the rulebook is written, you hand it to whichever agent you use. grok bot you just paste and go. hermes you deploy once on railway and talk to on telegram. either way, new work comes in, it gets checked, and you get the review list.
so your week basically looks like this. the agent pulls in each client's new work, runs it through the rulebook, and sends you one message per client. you look at the few things jev wasn't sure about, approve the batch, and it goes out.
you're running my agency. here's everything you need.
keys, stored as environment variables and never printed or logged:
- TYPESAFE_API_KEY for jev, from console.typesafe.ai
- ANTHROPIC_API_KEY for fable
- each client's stripe, shopify or accounting login, saved in that client's folder
how to call jev:
- POST https://api.typesafe.ai/v1/systemone with the model set to jev-latest
- send the case as the state, and that service's rulebook questions as the questions
- question types are choice for picking one option, score for a scale, and noul for true or false
- read back the answer and the probability for every question
how to call fable:
- anthropic api, model claude-fable-5-1
- only call it for cases jev cleared, to write the message, packet or letter
the loop:
- keep one folder per client and one rulebook per service
- weekly: pull each client's new invoices, disputes and bills from wherever they've connected
- run each case through its rulebook with jev
- anything over 0.8 gets written by fable and goes into a batch that's ready to send
- anything under 0.8 goes into a review list with the reason
- send me one message a week per client with the batch count, the review count and the money recovered so far
- never send, submit or dispute anything until i approve the batch
- if a client's data didn't come through or a key stops working, tell me rather than skipping them

where to start
pick one of the three services and one type of business. build three demo pages and send three cards or dms. the free first job is the sale.
hope you have fun setting this up.
join here for more value: https://t.me/+pbCBBtUEtu1lZDA1
Related articles

Grok Bot Agents: How to Build a 5-Person Sales Team That Never Sleeps (Full Guide)
I'm going to show you the exact 5-agent setup that's been running our lead gen, and the descriptions that go inside each one.

How to Build an AI-Native Company: Hiring, Standards, and Cadence
*We spent a year running hackathons, AI trainings, and office hours. But nothing changed until we made AI capability a requirement.*

HOW I WENT FROM $600/MONTH TO $14. ELON'S GROK + KIMI'S BRAIN
Follow & Bookmark this - I'm [@starmexxx](https://x.com/starmexxx), I track how AI tools are creating new income streams most people haven't heard of yet. This one is the entire build.