A support-ticket classifier running 10 million tokens a month, split roughly six input tokens for every one output token (long ticket threads and retrieved macros in, a short classification and reply out), is exactly the shape of workload where a new model’s headline input rate does most of the talking. Pareto landed on /models this week with that kind of headline. As of this writing, its listed input rate sits about 20% below GLM-4.5 Air’s, and its output rate sits about 15% above it.
At 10 million tokens with that 6:1 split, you’re paying for roughly 8.6 million input tokens and 1.4 million output tokens a month. Run that mix and Pareto comes out around 6% cheaper than GLM-4.5 Air for this exact pipeline, using GLM-4.5 Air’s current output-to-input rate multiplier (call it 4x, which is close to what that page shows today) to weight the two token types. The input discount survives because output is only 14% of total volume. That’s the workload where Pareto’s pricing does what the headline number promises.
Where the discount flips
Set up the algebra once and it stops mattering what the exact rates are next week. If GLM-4.5 Air’s output rate is k times its input rate, and Pareto’s input rate is 20% below GLM-4.5 Air’s while its output rate is 15% above, the two models cost the same when output tokens make up roughly 1/(1+4/k) of total volume. Plug in a 4x output multiplier (again, check GLM-4.5 Air’s live row for today’s actual figure since it moves) and the breakeven output share lands at 25%. Below that, Pareto wins. Above it, GLM-4.5 Air wins, and the gap widens fast as output share climbs.
Scale the same math to an agent loop at 100 million tokens a month with a 1:3 input-to-output ratio (25 million in, 75 million out, typical once tool calls and retries pile up output tokens), and output is now 75% of volume, three times past the breakeven point. Pareto comes out roughly 12% more expensive than GLM-4.5 Air on that mix, on the strength of the same output rate that only cost it 15% on paper. The input discount that looked decisive at 10 million tokens is now irrelevant. Output volume is doing all the work.
This is the trap in judging a new model off its input rate alone. A coding-agent team that sized a Pareto contract off a demo running near 1:1 input:output, then watched output climb toward 4:1 once retries and chain-of-thought reasoning kicked in on real tickets, would see monthly spend land somewhere around 30% over the number they budgeted, purely from the token mix shifting, with no change in volume or rate. Nobody re-quotes a contract because the ratio drifted. Run both ends of your actual ratio distribution through the calculator before committing, not just the demo mix.
What Pareto’s page doesn’t tell you yet
GLM-4.5 Air and MiniMax M2 both carry disclosed total and active parameter counts (106B/12B and 229B/10B respectively), which is what lets you size a self-hosted alternative and sanity-check API pricing against a GPU floor. Pareto’s model card doesn’t have that split published yet. No total-to-active ratio means no VRAM math, which means no way to tell whether the API rate is priced against thin provider margins on a small active-param model or wide margins on something bigger. Until that shows up, treat Pareto as API-only and don’t try to back into a self-hosting comparison from the rate alone. If you want a reference point for what an equivalent-class open model would cost to run yourself, GPUs still shows current floors for the hardware tier that class of model needs, but that’s a sanity check, not a quote for Pareto specifically.
If your workload’s output-heavy tail resembles the agent-loop example above, MiniMax M2.5 is worth pricing before either Pareto or GLM-4.5 Air. Its low active-param count is the kind of architecture that tends to keep output rates from ballooning the way a denser model’s does.
The decision rule
Pull last month’s actual token logs and compute your output share of total volume, not your input:output ratio from a demo run. If output tokens are under roughly a quarter of your total, Pareto’s input discount will hold and it beats GLM-4.5 Air on this workload. If output is a third or more of total volume, the output rate dominates and Pareto loses, sometimes by double digits, regardless of how good the input number looked on the page. Compute your own breakeven with the live multiplier on GLM-4.5 Air’s row before locking in either one.