ACTIVATED HUMAN/ ai

Which AI model should a business use?

Use the model that passes your own tests at the lowest cost your data rules allow, and expect to use more than one. For most businesses that means a large model for judgment calls, a smaller and cheaper one for sorting and extraction, and a second provider as a fallback during outages.

Why there is no single best model

Rankings change every few months, and the model at the top of a public leaderboard is often the most expensive one. Many routine tasks do not need it. Pulling fields from an invoice, sorting emails, or tagging a ticket are handled well by smaller models at a fraction of the price. Drafting a sensitive reply, reasoning over a contract, or planning a multi-step task may need the largest model available.

Data rules narrow the choice further. Some businesses cannot send client data to a provider that may retain it, and some need data to stay in a specific region or on their own hardware. So the question to answer is which model is good enough for each step of your work, at what cost, under your rules.

How we choose, task by task

QuestionHow we answer it
Is it accurate enough?Run each candidate model on your eval suite and compare pass rates on your cases.
What does it cost?Measure cost per run on real inputs, including cached and uncached calls.
Where does the data go?Check the provider's retention, training, and region terms against your policies.
Is it fast enough?Measure response time on real inputs, which matters most for anything a customer waits on.
What if it goes down?Pick a fallback provider that also passes the suite.

The candidates usually include Claude, OpenAI, Gemini, and open-source models. The result is written down per task, so you can see why each model was chosen and revisit the choice when prices or models change.

Running more than one provider

An agent that uses one provider stops working during that provider's outages. We run several model providers in production ourselves, with automatic fallback when one fails. For your systems, the fallback model has to pass the same eval suite as the main one.

For sensitive work we set a minimum model level for each task, and cost savings can never route the task below it. In one of our own products, safety-sensitive conversations always go to the most capable model, whatever the user's plan. An unknown or misspelled model name is treated as below every minimum, so a typo cannot quietly move sensitive work to a weaker model.

Open-source and self-hosted models

Open-source models make sense when data cannot leave your infrastructure, when volume is high enough that running your own hardware is cheaper, or when you need a model you control completely. They also bring work: hardware, updates, monitoring, and more careful evals, since smaller models vary more from task to task.

We can set up open-source models on your own servers or cloud account and run them beside hosted models, with each task routed to whichever passes its evals at the lowest cost. When self-hosting is an option, the audit compares its monthly cost with the hosted model's.

Getting this set up

If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.

Working with us
AI systems audit$6,500One week
We spend one week with the people doing the work.
Workflow agentsFrom $12,0002 to 4 weeks
An agent that takes over one recurring job and does it on its own in production, connected to your tools, with evals, tracing, and spending limits.
Agent systemsFrom $30,0004 to 8 weeks
Agent systems that run a core part of your business or your product in production, on your cloud or ours, with a custom harness, evals, tracing, and spending limits.

Questions

Is there a better AI model than ChatGPT for business?

For some tasks, yes. Claude, Gemini, and open-source models each lead on different kinds of work, and the right choice depends on your tasks, cost, and data rules. We test candidates on your own cases before recommending one.

Should a business use one AI model or several?

Usually several: a capable model for judgment calls, a cheaper one for routine steps, and a fallback provider for outages. Each is chosen and tested per task.

Which AI models keep business data private?

It depends on the provider, the plan, and the account settings, which control retention and whether data is used for training. We check those terms against your policies, and for strict requirements we can run open-source models on infrastructure you control.

Related questions