GPT-5.6 Sol vs. Terra vs. Luna: Did Paying More Change the Answer?

Read time: 8 minutes
GPT-5.6 Sol, Terra and Luna all choosing the marketing plan, with estimated costs of about $206, $69 and $8 per 1,000 similar analyses.

Written by: NisonCo Staff

Highlight your work with Public Relations

Find out how PR can support your marketing efforts.
Read more

We wanted to know whether paying for the most capable GPT-5.6 model would change a serious, numbers-based business decision.

What we tested

So we gave GPT-5.6 Sol, Terra and Luna the same fictional agency problem three times each. The agency had to choose between raising prices and risking client losses or keeping its current price and investing $40,000 in marketing.

What happened

Every memo chose the marketing plan, and the selected examples all gave the same correct base-case figures. Every answer arrived in under two minutes. If a business ran 1,000 similar analyses, the median estimates work out to about $8 with Luna, $69 with Terra and $206 with Sol.

The difference was reliability. One Luna memo had several wrong summary figures. A second Luna memo gave the correct recommendation, tables and required calculations but misstated one supporting cost figure.

The Short Answer

Every answer chose marketingSol, Terra and Luna all made the same recommendation.
Luna needed the closest checkingOne memo had several wrong totals. Another misstated one supporting cost figure.
Every answer took under two minutesIndividual runs finished in about 78 to 113 seconds.
About $8 to $206 for 1,000 similar analysesTerra would be about $69 on the same scale.

For this particular assignment, Terra offered the clearest balance. We found no errors in the supporting math we checked for either Terra or Sol, while Terra's estimated cost was much lower.

Why We Compared Sol, Terra and Luna

OpenAI offers different GPT-5.6 models for different combinations of capability, speed and cost. That creates a practical question for businesses: when does a lower-cost model become good enough for real work?

We did not want to answer that question with a generic prompt or a writing contest. We wanted a business assignment with a correct outcome that could be checked.

We used a fictional agency so each model could receive the same detailed information without exposing real financial data. We repeated each model three times because a single polished answer can hide inconsistency.

The Business Decision We Gave Each Model

The fictional agency had 20 clients paying $4,000 per month. It had two choices.

Option ARaise prices by 10%

Risk losing two clients, pay a $6,000 communication cost and slowly rebuild the client count.

Option BInvest $40,000 in marketing

Keep the $4,000 price, spend $10,000 per month for four months and add four clients on a supplied schedule.

Each memo had to calculate monthly revenue, operating profit and cash. It also had to test when the recommendation would change if the price increase retained more clients or the marketing plan attracted fewer clients.

Correct base-case answer: invest in marketing$395,000 in annual operating profit, which was $14,000 more than the price-increase plan.

What Sol, Terra and Luna Agreed On

We used the second completed memo from each model so the examples followed one fixed rule. All three selected memos gave the same correct answer: $395,000 in operating profit for marketing versus $381,000 for raising prices, a $14,000 difference.

GPT-5.6 SolAbout $206 for 1,000 similar analyses

“Option B is recommended under the stated base case. It produces $395,000 of 12-month operating profit, exceeding Option A’s $381,000 by $14,000.”

The recommendation and supporting calculations were consistent.

GPT-5.6 TerraAbout $69 for 1,000 similar analyses

“Recommend Option B under the stated base case. It delivers $14,000 more 12-month operating profit than Option A.”

The recommendation and supporting calculations were consistent.

GPT-5.6 LunaAbout $8 for 1,000 similar analyses

“Recommend Option B under the stated base case: it produces $395,000 of operating profit versus $381,000 for Option A, an advantage of $14,000.”

These main figures and its required tables were correct. Later in the same memo, one sentence misstated a supporting cost figure.

The decision and core figures shown here did not differ. The more expensive model did not uncover a different base-case answer.

Where Luna Needed More Checking

Every Luna memo chose the marketing plan, but two separate runs had different kinds of problems.

– One memo had several wrong summary figures. Its month-by-month rows were correct, but it reported the wrong annual profit, cash and downside totals.

Figure What the Luna memo said Correct figure
Marketing-plan annual operating profit $395,500 $395,000
Advantage over the price increase $14,500 $14,000
Profit if the first planned client never arrived $365,500 $365,000
Ending cash $545,500 $545,000
Lowest cash point $159,500 $169,500

The figures were close enough to look plausible. The memo even tried to explain one $500 discrepancy as a timing effect, although the discrepancy came from its own summary.

– The displayed Luna memo had the correct recommendation and main calculations but one bad explanation. It correctly calculated the $395,000 versus $381,000 result and the $14,000 advantage. It later said the difference in variable costs was $50,000. The correct difference was $16,000.

The correct reconciliation was $64,000 in additional revenue, minus $16,000 in additional variable costs, minus $34,000 in additional option-specific spending. That equals the correctly stated $14,000 advantage.

The second problem was smaller than the first, but it matters because the explanation did not support the right answer. It also explains why showing only the opening recommendation would make the models look more alike than their full memos were.

Three runs cannot establish a dependable failure rate. They can show that the cheapest generation price may require stronger automatic checks or more human review.

Two Luna runs chose marketing but needed different corrections: one had several wrong summary figures, and another had one incorrect supporting cost explanation.
The Luna runs agreed on the recommendation. One had several wrong summary and downside figures. Another had the correct recommendation and main calculations but one inaccurate cost explanation.

The Answer Stayed the Same. The Estimated Cost Did Not

All three models recommended marketing. The rows below show what changed: estimated cost and what we found when we checked the numbers.

Model and estimated cost for 1,000 similar analysesRelative costWhat our math check found
GPT-5.6 LunaAbout $8 for 1,000 similar analyses$0.0080 median per run
One memo had several wrong figures; another had one bad explanation
GPT-5.6 TerraAbout $69 for 1,000 similar analyses$0.0692 median per run
No errors in the supporting math
GPT-5.6 SolAbout $206 for 1,000 similar analyses$0.2055 median per run
No errors in the supporting math

These prices are estimates, not bills. We applied the OpenAI API prices recorded for the August 18 test to the token use reported by each run. The linked page may now show different rates. The per-1,000 figures simply multiply each model's median estimate by 1,000, so actual costs will change with the length of the work, token mix and current pricing. The experiment itself used an existing subscription, so it added $0 in cash spending.

Response times were much closer than costs. Every answer arrived in about 78 to 113 seconds. Model price did not change the base-case recommendation.

We found no errors in the supporting math we checked for Sol or Terra. One Luna memo had several wrong figures. A closer read found one wrong supporting figure in a different Luna memo whose recommendation and main calculations were correct.

Did Paying More for Sol Improve the Answer?

Not compared with Terra on this assignment.

Both models chose the marketing plan, and we found no errors in the supporting math we checked in any of their tested memos. At 1,000 similar analyses, their median estimates work out to about $206 for Sol and $69 for Terra, about 66% lower.

Terra was also faster in this small sample, with an 81.07-second median compared with Sol's 105.49 seconds.

That does not make Terra equal to Sol on every kind of work. A more complex task, missing information, research, tool use or a longer analysis could reveal a difference this assignment did not.

It does show that Sol's extra cost did not buy a better financial recommendation here.

Which GPT-5.6 Model Was Enough for This Job?

Terra was the strongest lower-cost starting point in this test.

It matched Sol's result in the supporting math we checked, arrived faster by median time and cost about one-third as much to generate based on reported token use.

Luna may still make sense for a high-volume first pass if the work is easy to verify automatically. Without a dependable check, its lowest token cost may be offset by the time required to find and correct both substantial mistakes and smaller inconsistencies.

The useful selection rule is not “always choose Terra” or “never choose Luna.” It is: choose the least expensive model that repeatedly handles the exact job your business needs.

How We Ran and Checked the Test

We tested GPT-5.6 Sol, Terra and Luna in Codex 0.145.0 with the medium effort setting on August 18, 2026.

Every model received the same fictional case and exact prompt. Each produced three first-pass memos without browsing, tools, follow-up instructions, retries or human corrections.

We checked every memo against a fixed answer key covering its recommendation, month-by-month values, annual totals, break-even calculations, downside cases and final answer.

GPT-5.5 also reviewed the memos in a blinded order for usefulness, reasoning and communication. Fable provided a second, exploratory review, but it was also one of the models in the full experiment, so we treated that review as a conflicted second opinion and did not use it to name a winner. Those secondary scores were not needed to determine whether the recommendation or calculations were correct. A planned separate human scorecard was not completed.

What This Test Does and Does Not Prove

This experiment shows what happened in nine runs of one fictional agency decision in Codex.

It does not prove that Terra is always equal to Sol or that Luna will make the same kind of mistake at a predictable rate. Three attempts per model are not enough to establish a general ranking.

The test did not cover web research, coding, tools, legal analysis, real financial advice, long-running work or decisions with missing information.

The cost figures are estimates based on reported token use and the published API prices we recorded on August 18, 2026. Vendor prices can change. The per-1,000 comparison is a simple extrapolation of each model's median per-run estimate, not a bill or volume quote. It does not include subscription allocation, employee review, corrections, setup or the business cost of a mistake.

The supported conclusion is narrower: all three models made the same recommendation, Terra matched Sol in the supporting math we checked at a much lower estimated price, and Luna's lowest cost came with more checking. One Luna memo had several wrong summary and downside figures, while another had the correct recommendation and main calculations but one inaccurate supporting figure.

The Bottom Line

We gave Sol, Terra and Luna the same business decision three times each to see whether paying more changed the answer.

It did not. All nine runs chose to keep prices steady and invest in marketing.

We found no errors in the supporting math we checked in the Sol or Terra memos. Terra produced that result at about one-third of Sol's median estimated cost. Luna cost far less again, but one memo used incorrect summary figures and another misstated one supporting cost figure.

For this assignment, Terra was already capable enough. Luna was only the better bargain if the workflow could catch its bad run.

For more ways to control model and workflow spending, read How to Reduce AI Costs Without Sacrificing Results.

For the full six-model result, read AI Model Comparison for Business Analysis: Our 18-Run Test.

Browse more NisonCo AI Tests & Comparisons.

Frequently Asked Questions

Did Sol, Terra and Luna choose the same strategy?

Yes. All nine runs recommended keeping the agency's current price and investing $40,000 in marketing.

Did all three GPT-5.6 models get the math right?

We found no errors in the supporting math we checked in the Sol or Terra memos. One Luna memo used incorrect summary figures even though its detailed monthly rows were correct. Another Luna memo had the correct recommendation and main calculations but misstated one supporting cost figure.

Was Terra cheaper than Sol?

Based on reported token use and published API prices, Terra's median estimated generation cost was $0.0692 per run compared with $0.2055 for Sol. That is about $69 versus $206 for 1,000 similar analyses, or about 66% lower.

Was Luna the cheapest model?

Yes. Luna's median estimated generation cost was $0.0080 per run, or about $8 for 1,000 similar analyses. Every Luna memo chose the correct strategy, but one used materially wrong summary figures and another misstated one supporting cost figure.

How long did the GPT-5.6 analyses take?

All nine runs finished in 77.97 to 113.18 seconds. Median times were 105.49 seconds for Sol, 81.07 seconds for Terra and 79.62 seconds for Luna.

Does this prove Terra is better than Sol?

No. It shows that Terra matched Sol on this one structured financial decision across three runs each. A different task could produce a different result.

Related posts

Skip to content