AI, Software & Automation
AI red-teaming and model-evaluation service
You are paid per engagement to attack a company's AI system before its customers do (jailbreaking the chatbot, poisoning the agent's tools, measuring where it fails) and hand back a scored report they can show a regulator.
- Advanced
- $5K–$25K
- Moderate risk
- 3–6 months to first customer
These bands place this model against the other 171 in the catalog so comparing them works — orientation, not a quote for your situation or your area. Figures that carry a source are on the Examples tab.
Why this stability rating: Regulation puts a floor under demand (the EU AI Act has required documented adversarial testing of general-purpose models with systemic risk since 2 August 2025) and every model release restarts the work, so engagements repeat. Against that: the buyer list is short, findings expire with the version they were found against, the automated half of the job is being productized into subscriptions, and a lab that hires four of your people can bring the work in-house.
- Asset-light
- Online
- Part-time friendly
- Inventory
Often fits: People comfortable learning technical tools, who enjoy solving one niche's problem deeply and can explain technology in the customer's language.
Often doesn't fit: People who want zero ongoing maintenance, hate keeping up with fast-moving tools, or want to avoid supporting clients when things break.
The simple explanation
Businesses everywhere pay people to do repetitive digital work: answering the same questions, moving data between systems, chasing leads, writing the same reports. This model replaces that work with software or AI, then charges for the result. You either build a product many customers use (SaaS) or install and maintain automations for specific clients (AI services). Either way, the thing you sell keeps working while you sleep. That is the leverage.
A simple hypothetical example
Illustrative — invented to show the shape of the AI & Software pattern. No real company is named, and no figure in it is data. The real, sourced companies for this model are on the Examples tab.
A local insurance broker types every new lead from their web form into three separate systems. You build an automation that does it instantly, charge a setup fee plus a monthly fee to keep it running, and the broker happily pays because it costs less than the hours it saves. Ten brokers later, you have recurring revenue and a repeatable playbook.
A closer look at ai red-teaming and model-evaluation service
You are selling a defensible negative result, which is a strange product: the client is paying you to fail to break something, and paying more when you succeed. Engagements price as fixed fees scoped by system and threat model (this chatbot, these tools, this data), and the recurring version is a retainer that re-runs the suite on every model version, prompt change or new tool the client wires in. The asset you build across clients is not the report; it is the harness. A reproducible attack suite that runs unattended is what turns a consultancy into something with margins, because the second client's sweep costs you inference rather than a week.
The market has split into three business shapes with very different economics, and picking one deliberately matters more than picking a niche. Automated continuous testing behaves like software: Gray Swan's $40 million Series A is funding a subscription whose marginal cost is tokens. Bespoke frontier evaluation behaves like a boutique: Irregular's simulations for OpenAI, Anthropic and Google DeepMind are sold on scarce expertise and bounded by how many senior people you can hire. In between sits the crowd, where you convene researchers and pay for findings. HackerOne reported record researcher payouts of more than $77.2 million in the year to March 2025, which is the cost line of that model laid bare.
Regulation is what makes this a business rather than a hobby, and it is also the ceiling on your price. Article 55 tells a lab it must document adversarial testing; it does not tell it to be impressed. A buyer procuring evidence wants an artefact, and artefacts commoditize. Within eighteen months there will be a template, a checklist and a competitor quoting a third of your fee for something that satisfies the same auditor. The escape is novelty: attacks nobody has published, an eval nobody else can run, a threat model specific to the client's actual deployment rather than to models in general. Sell the finding, not the PDF.
Two things quietly erode this business from underneath. The first is expiry. Your findings are pinned to a version, and a model update can retire an entire report before the invoice clears, which is exactly why the retainer is the honest sale and the one-off audit is not. The second is the conflict problem, which almost nobody plans for. Once your evaluations are cited in competing labs' system cards, every client is entitled to ask what you learned inside their rival, and your answer has to be a written information wall rather than a reassuring sentence. Add to that the legal floor: never touch a system without signed authorization and a scope that names the systems, the window and the data you may retain. Enthusiastic unauthorized testing is not a demonstration of value, it is a computer-misuse problem with your name on it.
How money moves through this model
Who pays: Businesses (usually) or consumers paying a subscription or setup + retainer
What they pay for: Time saved, errors avoided, or capability they can't build themselves
What creates profit: The gap between what the automation earns you monthly and the small cost of running it
- Customer
- Offer
- AI
- Costs
- Profit
What makes this model hard
The honest difficulty: the technology is the easy half. The actual business is finding a niche where the same automation sells over and over, explaining it to non-technical buyers, and supporting it when it breaks at 9pm. Tools change fast, and what feels like a moat today can become a commodity feature next year.