Six AI Models, One Business Decision: Did Paying More Change the Answer?

Read time: 10 minutes
Six AI models comparing a price increase with a marketing investment, with estimated costs ranging from about $8 to $455 per 1,000 similar analyses.

Written by: NisonCo Staff

Highlight your work with Public Relations

Find out how PR can support your marketing efforts.
Read more

We wanted to know whether paying for a more expensive AI model would improve a serious, numbers-based business analysis.

What we tested

To test that assumption, we gave six AI models the same fictional agency problem. The agency had to choose between raising its prices and risking client losses or keeping its current price and investing $40,000 in marketing. We ran each model three times, producing 18 separate decision memos.

What happened

The result was unusually clear. Every memo chose the same strategy: keep the current price and invest in marketing. Seventeen of the 18 memos matched our answer key. A closer read found one separate inaccurate explanation in another Luna memo whose recommendation and main calculations were correct. Every answer arrived in under two minutes. If a business ran 1,000 similar analyses, the median estimates work out to about $8 with the least expensive model and $455 with the most expensive one.

The Short Answer

Every answer chose marketingAll six models made the same business recommendation.
Seventeen of 18 memos matched the answer keyA closer read found one separate bad explanation in another Luna memo.
Every answer took under two minutesIndividual runs finished in about 66 to 113 seconds.
About $8 to $455 for 1,000 similar analysesThe median estimated cost varied by about 57 times.

For this task, several lower-cost models were already capable enough. The cheapest option was only truly cheaper if the workflow included a reliable way to catch both substantial mistakes and smaller inconsistencies.

Why We Ran This Test

Businesses now have a growing menu of AI models with different prices and capability claims. The default reaction is often to choose the most powerful model for anything important.

That can be sensible for difficult work. It can also mean paying more for capability a particular task does not need.

We wanted to test a narrower and more useful question: if a business decision has clear facts and calculations that can be checked, will lower-cost models reach and support the same answer as more expensive ones?

We used a fictional agency so every model could receive exactly the same information without exposing a real company's financial data. We repeated the assignment three times per model because one polished answer can hide inconsistency.

The six tested models were GPT-5.6 Sol, Terra and Luna in Codex, plus Claude Sonnet 5, Opus 4.8 and Fable 5 in Claude Code.

The Business Decision We Gave Every Model

The fictional agency started with 20 clients paying $4,000 per month. It had two choices.

Option ARaise prices by 10%

Risk losing two clients, pay a $6,000 communication cost and slowly rebuild the client count.

Option BInvest $40,000 in marketing

Keep the $4,000 price, spend $10,000 per month for four months and add four clients on a supplied schedule.

The models had to calculate the agency's revenue, operating profit and cash for every month. They also had to test what would happen if the price increase retained an extra client or the marketing plan brought in fewer new clients than expected.

Correct base-case answer: invest in marketing$395,000 in annual operating profit, which was $14,000 more than the price-increase plan.

Option B was the correct recommendation under the supplied assumptions, but its advantage was not enormous. That made the break-even and downside calculations important.

The Answer Stayed the Same. The Estimated Cost Changed a Lot

Every model recommended the marketing plan. What changed was the estimated price and how closely the outputs needed to be checked.

Model and estimated cost for 1,000 similar analysesRelative costWhat our math check found
GPT-5.6 LunaAbout $8 for 1,000 similar analyses$0.0080 median per run
One memo had several wrong figures; another had one bad explanation
GPT-5.6 TerraAbout $69 for 1,000 similar analyses$0.0692 median per run
No errors in the supporting math
Claude Sonnet 5About $118 for 1,000 similar analyses$0.1181 median per run
No errors in the supporting math
GPT-5.6 SolAbout $206 for 1,000 similar analyses$0.2055 median per run
No errors in the supporting math
Claude Opus 4.8About $248 for 1,000 similar analyses$0.2479 median per run
No errors in the supporting math
Claude Fable 5About $455 for 1,000 similar analyses$0.4552 median per run
No errors in the supporting math

The cost estimate applies the OpenAI and Anthropic API prices recorded for the August 18 test to the token use reported by each run. These figures are estimates, not bills, and the linked pages may now show different rates. The per-1,000 figures simply multiply each model's median estimate by 1,000, so actual costs will change with the length of the work, token mix and current pricing. The experiment itself used existing subscriptions, so it added $0 in cash spending.

Response times were much closer than costs. Every answer arrived in about 66 to 113 seconds. The most useful result is that the lower-cost models reached the same decision as the premium models on a task that required business judgment, monthly projections, break-even math and downside analysis.

What the Six Models Agreed On

To avoid choosing a flattering example for one model and a weak example for another, we used the second completed memo from each model below. All six selected examples gave the same correct base-case answer. That agreement is the point of this section.

GPT-5.6 SolAbout $206 for 1,000 similar analyses

“Option B is recommended under the stated base case. It produces $395,000 of 12-month operating profit, exceeding Option A’s $381,000 by $14,000.”

The recommendation and supporting calculations were consistent.

GPT-5.6 TerraAbout $69 for 1,000 similar analyses

“Recommend Option B under the stated base case. It delivers $14,000 more 12-month operating profit than Option A.”

The recommendation and supporting calculations were consistent.

GPT-5.6 LunaAbout $8 for 1,000 similar analyses

“Recommend Option B under the stated base case: it produces $395,000 of operating profit versus $381,000 for Option A, an advantage of $14,000.”

These main figures and its required tables were correct. Later in the same memo, one sentence misstated a supporting cost figure.

Claude Sonnet 5About $118 for 1,000 similar analyses

“Under the stated base case, Option B, hold price and invest in marketing, is recommended over Option A.”

Sonnet immediately gave the same $395,000 versus $381,000 comparison.

Claude Opus 4.8About $248 for 1,000 similar analyses

“Adopt Option B, hold price and invest in marketing.”

Opus emphasized that the $14,000 lead was thin and execution could change the result.

Claude Fable 5About $455 for 1,000 similar analyses

“Recommend Option B, hold price and invest in marketing.”

Fable gave the same profit totals and pointed to the supplied downside cases.

Readers can judge those answers directly. They are not identical, but none of the more expensive models found a different base-case decision or a better profit total. The reliability differences appeared elsewhere in the supporting work.

Where Luna Needed More Checking

The least expensive model, GPT-5.6 Luna, recommended Option B every time. Two separate runs had different kinds of problems.

– One memo had several wrong summary figures. Its detailed monthly rows were correct, but its summary did not agree with them.

Figure What the Luna memo said Correct figure
Option B annual operating profit $395,500 $395,000
Option B advantage over Option A $14,500 $14,000
Option B profit if its first planned client never arrived $365,500 $365,000
Option B ending cash $545,500 $545,000
Option B lowest cash point $159,500 $169,500

The memo even tried to explain one $500 discrepancy as a timing effect. The discrepancy came from its own summary.

– The displayed Luna memo had the correct recommendation and main calculations but one bad explanation. It correctly calculated the $395,000 versus $381,000 result and the $14,000 advantage. It later said the difference in variable costs was $50,000. The correct difference was $16,000.

The correct reconciliation was $64,000 in additional revenue, minus $16,000 in additional variable costs, minus $34,000 in additional option-specific spending. That equals the correctly stated $14,000 advantage.

Both kinds of problems can survive a quick read. The first memo's wrong figures were caught by our answer-key check. The second required a closer read because its recommendation and main calculations were right.

Three runs are far too few to claim that Luna has a predictable failure rate. They are enough to show why “cheapest” and “lowest total cost” are not always the same thing. If a person has to recheck every figure manually, the savings can disappear. If the workflow automatically compares the memo with a trusted calculation and reviews its explanation, the problems are much easier to catch.

Two Luna runs chose marketing but needed different corrections: one had several wrong summary figures, and another had one incorrect supporting cost explanation.
The Luna runs agreed on the recommendation. One had several wrong summary and downside figures. Another had the correct recommendation and main calculations but one inaccurate cost explanation.

Did Paying More Produce a Better Answer?

Not in the outcome we tested.

We found no errors in the supporting math we checked in any tested memo from GPT-5.6 Sol, GPT-5.6 Terra, Claude Sonnet 5, Claude Opus 4.8 or Claude Fable 5. At 1,000 similar analyses, their median estimates still ranged from about $69 to $455.

The clearest comparisons happened within each product.

– In Codex: Terra matched Sol in the supporting math we checked at about one-third of Sol's median estimated cost. Terra was also about 24 seconds faster by median runtime.

– In Claude Code: We found no errors in the supporting math we checked in the Sonnet, Opus or Fable memos. Sonnet's median estimate was about half of Opus's and about one-quarter of Fable's.

Fable was the fastest model in this test. Its typical run was about 13 seconds faster than Sonnet's. For a single memo, that time saving probably would not justify paying almost four times as much for generation. A long-running agent workflow could produce a different tradeoff, and this experiment did not test that kind of work.

The practical conclusion is simple: more expensive did not mean more correct for this particular assignment.

Which AI Models Were Enough for This Business Analysis?

Do not start by asking which model is best overall. Start with a real task your team performs repeatedly.

– Define what a correct answer looks like. For this test, that meant the right recommendation, exact monthly numbers, break-even calculations and downside cases.

– Try more than one model. Include a lower-cost option and a premium option instead of assuming the premium one is necessary.

– Run each model more than once. A single polished response can hide an occasional mistake.

– Count review and correction time. A cheaper model is not cheaper if employees spend much longer fixing its work.

– Pay more for a measured reason. Move to a premium model when it reduces errors, saves meaningful review time or provides a capability the lower-cost model lacks.

For this fictional financial memo, Terra and Sonnet were the clearest lower-cost starting points. We found no errors in the supporting math we checked in their tested memos, and both cost much less to generate than the more expensive model in the same product.

Your workflow may point somewhere else. The useful goal is not to choose the cheapest model. It is to find the least expensive model that repeatedly produces work your business can accept.

Five-step AI model selection process: define the task, set checks, repeat runs, count review and correction, and pay more only for a measured reason.
Start with a real task and the least expensive model that repeatedly handles it correctly. Pay more when a measured failure, review burden or missing capability justifies it.

How We Ran and Checked the Test

We tested GPT-5.6 Sol, Terra and Luna in Codex 0.145.0. We tested Claude Sonnet 5, Opus 4.8 and Fable 5 in Claude Code 2.1.220. Each model used the product's medium effort setting on August 18, 2026.

Every model received the same fictional case and exact prompt. Each produced three first-pass memos without browsing, tools, follow-up instructions, retries or human corrections.

We compared each memo's monthly values, annual totals, recommendation, break-even calculations, downside cases and final answer with a fixed answer key. Seventeen memos matched every value in that automated check. On a closer editorial read, we found the separate explanatory error described above in one of those 17.

GPT-5.5 also reviewed the memos in a blinded order for usefulness, reasoning and communication. Fable provided a second, exploratory review, but it was also one of the tested models, so we treated that review as a conflicted second opinion and did not use it to name a winner. Those secondary scores were not needed to determine whether the recommendations or calculations were correct. A planned separate human scorecard was not completed.

What This Test Does and Does Not Prove

This experiment shows what happened in 18 runs of one bounded fictional agency decision.

It does not prove that one model is always better or that the six models used equal computing effort. Codex and Claude Code have different system instructions and execution environments. A “medium” setting in one product does not guarantee the same amount of work in another.

The test also did not cover research, web browsing, coding, legal analysis, real financial advice, long-running agents or decisions with missing information. Any of those could produce different results.

The cost figures are estimates based on reported token use and the published API prices we recorded on August 18, 2026. Vendor prices can change. The per-1,000 comparison is a simple extrapolation of each model's median per-run estimate, not a bill or volume quote. It does not include subscription allocation, employee review, corrections, setup or the business cost of an error.

The safe conclusion is narrower: for this structured, checkable assignment, we found no errors in the supporting math we checked in any memo from five models despite large differences in estimated cost. The cheapest model chose the same strategy every time, but one memo had several wrong summary and downside figures and another had one inaccurate supporting explanation.

The Bottom Line

We gave six AI models the same business decision three times each because we wanted to know whether paying more would change the answer.

It did not. Every run chose to keep prices steady and invest in marketing. Seventeen of the 18 memos matched our answer key, and every answer arrived in under two minutes. A closer read found one separate inaccurate supporting explanation in another Luna memo whose recommendation and main calculations were correct.

Several lower-cost models were already capable enough for the job. The memo with several wrong figures and the additional bad explanation both came from the cheapest option, which is why a good AI workflow needs both a capable model and a reliable way to check its work.

Before paying for the most expensive model, test a less expensive one on the exact work you need. Then pay more only when the results give you a clear reason.

For more ways to control model and workflow spending, read How to Reduce AI Costs Without Sacrificing Results.

For the vendor-specific results, read GPT-5.6 Sol vs. Terra vs. Luna and Claude Sonnet vs. Opus vs. Fable.

Browse more NisonCo AI Tests & Comparisons.

Frequently Asked Questions

What business decision did the six AI models analyze?

They analyzed whether a fictional agency should raise its monthly price and risk losing clients or keep its current price and spend $40,000 on marketing.

Did all six AI models choose the same strategy?

Yes. All 18 runs recommended keeping the current price and investing in marketing.

Did every model get the calculations right?

Seventeen of the 18 memos matched the answer key in our automated check. One GPT-5.6 Luna memo chose the right strategy but used materially wrong figures in its summary and part of its downside analysis. Another Luna memo had the correct recommendation and main calculations but misstated one supporting cost figure.

Were the more expensive AI models better?

Not in this test. Sol, Terra, Sonnet, Opus and Fable made the same decision, and we found no errors in the supporting math we checked in their memos, despite large differences in estimated cost.

How long did the AI business analyses take?

Every run finished in 65.81 to 113.18 seconds. Median times by model ranged from 66.22 seconds for Fable to 105.49 seconds for Sol.

How much did the AI analyses cost?

The experiment used existing subscriptions, so it added $0 in cash spending. Based on reported token use and published API prices, median estimated generation cost ranged from $0.0080 per run for Luna to $0.4552 for Fable. That is about $8 versus $455 for 1,000 similar analyses.

Does this prove which AI model is best for business analysis?

No. It compares six models on one fictional, structured decision with three runs each. A different task could produce a different result.

Related posts

Skip to content