GPT-Astra, Sol, Terra, and Luna: four new tiers, priced side by side

Published 2026-10-07
Four same-week model tiers priced daily on llmhosting.ai

A team migrating their coding agent last week picked Terra over Luna because Terra sounded like the bigger, more capable name. Two weeks and a few million tokens later, their monthly bill came in noticeably above what Luna would have cost for the same completions. Terra’s input rate ran well above Luna’s, and their agent workload is input-heavy: full file context in, a small diff out. The name told them nothing. The row on /models would have.

Four launches, zero naming signal

When four models ship within a day of each other from the same lab, the names get picked for branding, not for a consistent price or capability ladder. Astra does not mean “entry tier.” Luna does not mean “cheap.” Sol is not necessarily the fast one. Each of the four carries its own input rate, output rate, context window, and latency tier, and none of those four variables sort the same way across the set. The only way to know which is cheapest for your traffic is to pull all four rows and build the table yourself.

That table needs five columns: input $/M, output $/M, context window, latency tier (standard versus a priority or “fast” SKU that often carries a token premium), and whatever reasoning or thinking mode the model exposes, since some of these launches bill reasoning tokens separately from output. Skip any one of those columns and you’ll rank the four wrong for a workload that doesn’t match the demo traffic the lab benchmarked against.

Why coding agents break the obvious ranking

Astrahigh in / low out Solmid in / high out Terralow in / mid out Lunalowest in / mid out blended $/M onlanding page (1:1 ratio) your real ratio:input-heavy agent(full file in, small diff out) Luna wins onyour traffic re-rank byyour ratio name ladder ≠ price ladder — pull all 4 rows yourself Names don't sort price — input weight does

A chat workload and a coding-agent workload have almost opposite token shapes. Chat skews toward balanced or output-heavy exchanges. A coding agent reading a repo, running a tool, and producing a patch skews hard toward input: full file contents, test output, prior turns, all going in, with a comparatively small diff coming out. So a model with a slightly higher output rate but a meaningfully lower input rate can beat a model that looks cheaper on a blended “average $/M” marketing number.

That’s the actual trap in the Astra/Sol/Terra/Luna set. Whichever one publishes the lowest blended rate on a landing page is not automatically the cheapest for you, because blended rates assume a 1:1 or 1:3 ratio that your traffic probably doesn’t match. Pull your own input-to-output ratio from last month’s logs, not from a vendor’s example prompt, and apply it to each of the four rows.

Context window changes the math a second way. If Terra’s window is smaller than the largest file set your agent regularly loads, you’re not comparing Terra’s nominal rate against Luna’s; you’re comparing Terra’s rate plus the overhead of chunking a repo into multiple calls against Luna’s single-pass rate. Chunking adds repeated system-prompt and header tokens on every chunk. A model that looks 15% cheaper per token can lose that whole margin to a 20% token inflation from chunking, and the only way to catch this before it shows up on an invoice is to check each model’s context window against your real file sizes, not against the lab’s stated maximum.

Running the actual spread

smaller window(~15% cheaper/tok) repo too big→ split into chunks repeated systemprompt + headers(~20% inflation) net: marginerased single-passwindow fits repo 50M-tok monththrough calculator compare vs openGPT-OSS 120B onH100 SXM floor no chunking sanity check check window vs your real file sizes, not the lab's stated max A 'cheaper' context window can cost more

Once you have all four rows, run a matched 50M-token month through /calculator using your real input:output ratio for each candidate. Fifty million tokens a month is a realistic mid-size coding-agent fleet: a few dozen engineers running agentic sessions daily, each session reading meaningful context and emitting patches, tool calls, and commit messages. At that volume the spread between the cheapest and priciest of the four isn’t a rounding error. It’s the kind of gap that funds or kills a second hire. Provider-level spreads on the same open model already run past 20x on this site for models like qwen3-235b-a22b and deepseek-v4-flash as of this writing, per the live rows on their model pages; a same-lab four-tier spread on new proprietary launches is routinely smaller than that, but it is never zero, and name-based guessing gets it wrong often enough to matter.

If your monthly volume at 50M tokens pushes the calculator’s output anywhere near what a dedicated GPU would cost you, check whether an open-weight alternative closes the gap. GPT-OSS 120B runs 5.1B active params against 117B total, which at the roughly 1.1GB-per-billion-param FP8 rule plus 25% KV headroom fits comfortably on a single H100 SXM, and its per-token API rate on /models gives you a real floor to compare Astra, Sol, Terra, and Luna against before you assume a proprietary tier is your only option.

What this can’t tell you

I can’t confirm whether Astra, Sol, Terra, and Luna actually sit on a consistent capability ladder or whether they’re four independently tuned models that happen to share a launch week. If you downgrade your agent to whichever tier is cheapest on paper and your patch quality drops, your retry count eats the savings fast, sometimes within the first week. Before committing a fleet to any of the four, run a twenty-session matched sample across all four tiers on your real repo and track both cost per completed task and retries per task, not just $/M. Check /trends in a month to see whether the spread between these four held, widened, or collapsed once the initial launch pricing settles.

Put the numbers to work. Run your own workload through the breakeven calculator with today's floors, then price hardware: rent at Vast.ai → · rent at RunPod → Rental links are referral links. They never affect our rankings, which are ordered by price alone.

← All guides