Sign in

fronset.ai

@fronset.ai
61 followers 11 following 1.2K posts

The cheapest model good enough, per call. An OpenAI-compatible API that routes each request to the right-sized LLM for the task — backed by our independent benchmark on real production workloads. fronset.ai

PostsRepliesMedia
fronset.ai @fronset.ai · 07/10/2026
If your planner routes on the published 24-hour batch window, it silently overvalues the latency edge of pricier modes. With 76.0% of 57,219 completed batches back inside 15 minutes, the measured distribution — not the SLA text — should pick the channel.
000
fronset.ai @fronset.ai · 07/10/2026
First-pass reliability is not uniform across models. meta/muse-spark-1.3-contributor returns a usable answer first time on 99.9% of 24,572 attempts, while GPT-5.6 Luna hits 98.7% of 139,496 attempts and Gemini 3.5 Flash reaches 96.2% of 93,356 attempts.
000
fronset.ai @fronset.ai · 07/10/2026
If you route on quality alone, you miss reliability: Claude Sonnet 5's 8.1 out of 10 on Structured Content Summarization describes only the 85.7% of 102 attempts that answered. Read quality and answer rate as two numbers.
000
fronset.ai @fronset.ai · 07/10/2026
If Gemini 3.1 Flash Lite scores 5.6/10 only when it answers, which happens on 88.6% of 266 attempts, doesn’t that mean the headline quality score overstates delivered performance without its answer rate attached?
000
fronset.ai @fronset.ai · 07/10/2026
Cheap does not mean sloppy: GPT-5.6 Luna delivers a usable first-time answer on 98.7% of 139,496 attempts; Gemini 3.1 Flash Lite on 97.1% of 196,307; Gemini 3.5 Flash on 96.2% of 93,356; meta/muse-spark-1.3-contributor on 99.9% of 24,572—no retry required.
000
fronset.ai @fronset.ai · 06/10/2026
Published provider SLA for batch: up to 24 hours. Observed across 57,219 completed batches: median 5.4 minutes, 76.0% done inside 15 minutes. The gap between the contract and the behaviour is where the savings on summarization work are sitting.
000
fronset.ai @fronset.ai · 06/10/2026
If you rank models on quality alone, you rank them on their successes. NVIDIA Nemotron 3.5 Lightning's 8.08/10 on Social Post Portfolio Selection sits on top of a 89.7% success rate (n=34) — roughly one attempt in ten never produced an answer to score.
000
fronset.ai @fronset.ai · 06/10/2026
On the content-domain suggestion task, meta/muse-spark-1.1 was judged the best answer in 39 of the 45 occasions it competed in — 86.7% of appearances (n=45). Every win came head-to-head against other models answering the same live input.
000
fronset.ai @fronset.ai · 06/10/2026
Some tasks reward model choice: on at_content_domain_suggest, meta/muse-spark-1.1 was judged best in 39 of 45 occasions (86.7%). Language detection is the opposite — top two answers tied in 97.2% of 71 judged occasions, median gap 0.0 out of 10.
000
fronset.ai @fronset.ai · 06/10/2026
If a model scores 5.6 out of 10 on a task but returns nothing on roughly 1 in 9 attempts (88.6% of 266 for Gemini 3.1 Flash Lite on Known Url Content Extraction), what is that score actually describing? Answered calls only. Check the measurement flags before you trust it.
000
fronset.ai @fronset.ai · 06/10/2026
Same model, same task, two price tags: GPT-5.4 Nano on Factual Claim Refinement costs $0.000422 per task via batch versus $0.001135 synchronously. That's 62.8% off the sync cost, purely for agreeing to wait.
000
fronset.ai @fronset.ai · 06/10/2026
Four tasks, same pattern: billed tokens that couldn't be parsed took 6.3% of spend on relevance_scoring_post, 6.8% on author_matching, 6.1% on bench_editorial_story_planning and 9.7% on engagement_review. A single-digit tax that never shows up on the price list.
000
fronset.ai @fronset.ai · 05/10/2026
Before trusting a 5.6 out of 10, ask: 5.6 on which attempts? Gemini 3.1 Flash Lite answers only 88.6% of 266 calls on Known Url Content Extraction — about 11% of calls return nothing usable. What's the score worth if the answer rate isn't priced in?
000
fronset.ai @fronset.ai · 05/10/2026
For classification or matching work that can wait, choose your cost mode before your model. Running GPT-5.4 Nano on Factual Claim Refinement in batch costs $0.0004 per task versus $0.0011 synchronously — 62.8% less than the sync price, before any model swap is even considered.
000
fronset.ai @fronset.ai · 05/10/2026
Point scores on relevance and classification work barely move: top two answers identical on 97.2% of 71 judged language-detection occasions. Costs move a lot: GPT-5.4 Nano on Factual Claim Refinement runs $0.0004 per task in batch vs $0.0011 synchronously.
000
fronset.ai @fronset.ai · 05/10/2026
Implication: using one global timeout would either cut off prompt‑adaptation (31 s median, n=11,453) or waste resources on language‑detection (1.8 s median, n=69,136). Task‑specific limits avoid both.
000
fronset.ai @fronset.ai · 05/10/2026
On language-detection, DeepSeek V4 Flash’s median latency is 1.8 s (69,136 calls) vs qwen3.7‑flash’s 6.0 s (10,521 calls); thus deploying a different model variant shifts median latency.
000
fronset.ai @fronset.ai · 05/10/2026
Contrast shows task matters: DeepSeek V4 Flash finishes language‑detection in 1.8 s (69,136 calls) but needs 16 s for claim_extraction (56,798 calls) under the same 600‑second deadline.
000
fronset.ai @fronset.ai · 05/10/2026
Because Gemini 3.1 Flash Lite’s quality score conditions on successful calls only, relying on the headline 5.6‑point figure overstates true usefulness; model comparisons should adjust for the 88.6 % answer rate (n=266) to avoid overestimating performance.
000
fronset.ai @fronset.ai · 05/10/2026
Unparseable output drains 6–10% of task spend across jobs. On relevance_scoring_post, $29.01 of $462.25 went to unreadable tokens. On author_matching, $3.77 of $55.41. On engagement_review, $3.02 of $31.11. The cost is real and must be priced in.
000
fronset.ai @fronset.ai · 04/10/2026
GPT-5.4 Nano's sync price on Factual Claim Refinement is $0.001135 per task. The same task through batch: $0.000422, a 62.8% saving (n=1). Same work, same model, one difference — whether it can wait.
000
fronset.ai @fronset.ai · 04/10/2026
Caveat: averages reflect only returned calls. DeepSeek V4 Flash scores 7.43/10 on Atomic Fact Claim Extraction when it answers, but answers on 12.3% of 430 attempts. Means miss non-response — and never surfaced DeepSeek V4 Pro’s 0 wins in 91 judged claim_refinement occasions.
000
fronset.ai @fronset.ai · 04/10/2026
Claim: Conditional quality does not transfer across neighboring tasks. Qwen 3.7 Plus scores 9.67/10 (31 judgments) on Structured Output Extraction but only 5.87/10 on Markdown Newline Repair, answering just 27.9% of 46 attempts...
000
fronset.ai @fronset.ai · 04/10/2026
Two numbers, same scale. Qwen 3.7 Plus: 9.67 out of 10 on Structured Output Extraction after 31 independent judgments. DeepSeek V4 Flash: 7.43 out of 10 on Atomic Fact Claim Extraction — on the 12.3% of 430 attempts it answered at all. Only one tells you what ships.
000
fronset.ai @fronset.ai · 04/10/2026
This is a bias in our own scoring. claim_refinement winners ran 1,561 output tokens (n=143) vs 1,762 for losers (n=2076); claim_extraction winners ran 2,163 (n=156) vs 1,925 (n=2055). The favored direction flips by task — no global brevity setting holds.
000
fronset.ai @fronset.ai · 04/10/2026
On Catalyst And Scenario Analysis, Claude Opus 5 was judged the best answer in 10 of the 30 occasions it competed in — 33.3%, each against other models answering the same live input. A plurality lead, not a takeover: it loses two contests in three.
000
fronset.ai @fronset.ai · 03/10/2026
When NVIDIA Nemotron-3 Super 120B returns usable answers first time on 95.2% of 2,466 attempts at scale, it proves near-complete first-time delivery is achievable. We should measure and compare this across all models.
000
fronset.ai @fronset.ai · 03/10/2026
A model’s headline score can look strong while missing many tries. Nemotron 3.5 Lightning scores 8.08/10 on Social Post Portfolio Selection when it answers, but it answers only 89.7 % of its 34 attempts. Without the answer‑rate, the score overstates what a pipeline actually receives.
010
fronset.ai @fronset.ai · 03/10/2026
On language-detection, the top two models scored identically in 97.2% of 71 judged occasions with a median gap of 0.0 points. Task choice barely moved results—so relying on cross-task reputation instead of task-specific win records makes no sense here.
000
fronset.ai @fronset.ai · 03/10/2026
The caveat is clear: Gemini 3.1 Flash Lite’s 5.6/10 score applies only to the 88.6% of attempts that returned a response (n=266). Without adjusting for the missing answers, the score overstates reliability.
000
fronset.ai @fronset.ai · 03/10/2026
Batch mode sharply cuts per‑run cost: GPT‑5.4 Nano on Factual Claim Refinement runs at $0.000422 per task in batch vs $0.001135 sync, saving 62.8% of sync cost (n=1).
000
fronset.ai @fronset.ai · 03/10/2026
If output is unparsable, the token cost remains real. Engagement_review spent $3.02 of $31.11 (9.7% of task spend, n=35) on billed but unparsable output. This hidden tax means the effective cost per usable result is significantly higher than the raw billing suggests.
000
fronset.ai @fronset.ai · 03/10/2026
Qwen 3.7 Plus scores 5.87/10 on Markdown Newline Repair AND only answers 27.9% of 46 attempts — weak quality and weak coverage together. Without retries or alternate paths, most failures here will be silence, not wrong answers.
000
fronset.ai @fronset.ai · 03/10/2026
A tight confidence interval only measures precision on answered calls, not coverage. Gemini 3.5 Flash posts 9.33 out of 10 (±0.14 points after 39 independent judgments) on Markdown Newline Repair, but that narrow band says nothing about attempts that failed to return.
000
fronset.ai @fronset.ai · 02/10/2026
DeepSeek V4 Flash looks solid at 7.43 out of 10 on Atomic Fact Claim Extraction — until you add coverage: 12.3% of attempts (n=430). Answered-only averages hide the misses; deployable performance has to count non-answers too.
000
fronset.ai @fronset.ai · 02/10/2026
If we rank models solely by judged scores, we will pick verbose models for content-summarization (3,333 vs 2,203 tokens, n=146/1,953) and engagement_triage (744 vs 618 tokens, n=103/1,222) but concise models for claim_refinement (1,561 vs 1,762 tokens, n=143/2,076).
000
fronset.ai @fronset.ai · 02/10/2026
Claude Opus 5 secured the top spot in 33.3% of its 30 appearances on Catalyst And Scenario Analysis. While an impressive win rate, it underscores how competitive the field remains when models face the same live inputs.
000
fronset.ai @fronset.ai · 02/10/2026
On claim_refinement judges picked shorter answers: 1,561 output tokens on winning calls (n=143) vs 1,762 on losing calls (n=2076). On claim_extraction they picked longer: 2,163 (n=156) vs 1,925 (n=2055). Neighboring tasks, opposite length preference.
001
fronset.ai @fronset.ai · 02/10/2026
Qwen 3.7 Plus sets the Structured Output Extraction bar at 9.67 out of 10, ±0.15 over 31 judgments — but that's quality when it answers. Compare: Gemini 3.1 Flash Lite scores 7.5 out of 10 on teaser generation while returning on 0% of 55 attempts. Coverage is a separate test.
000
fronset.ai @fronset.ai · 02/10/2026
On Catalyst And Scenario Analysis, the top two answers tied exactly in 18.2% of 33 judged occasions and landed within half a point in 66.7% — median gap 0.2 on a 10-point scale. Score can't settle it; coverage should.
000
fronset.ai @fronset.ai · 02/10/2026
Two numbers, same model: 7.5 out of 10 quality on Short Promotional Teaser Generation, 0% answer rate across 55 attempts. Conditional quality says "good enough". Answer rate says the work never gets done. Route on the second one.
000
fronset.ai @fronset.ai · 01/10/2026
Because serving mode changes cost while model and task stay fixed, any per-run price quoted for a step is only meaningful if you state whether it was batch or sync. For GPT-5.4 Nano on Factual Claim Refinement, batch is $0.000422 per task and sync is $0.001135 per task (n=1).
000
fronset.ai @fronset.ai · 01/10/2026
Step-level routing must be priced on usable output, not nominal per‑run cost, because 6.1%‑9.7% of spend buys unparseable tokens: $0.81 of $13.34 on bench_editorial_story_planning (n=23), $3.02 of $31.11 on engagement_review (n=35), $3.77 of $55.41 on author_matching (n=38), $29.01 of $462.25 on...
000
fronset.ai @fronset.ai · 01/10/2026
Looking at the 5.6/10 score without the 88.6% success rate (n=266) inflates perceived performance; together they show the true expected quality across all attempts.
000
fronset.ai @fronset.ai · 01/10/2026
If answer rate is under 100%, the headline quality score overstates what you'd see in production. Gemini 3.1 Flash Lite: 5.6 out of 10 on Known Url Content Extraction, measured on the 88.6% of 266 attempts that returned something. The rest need a fallback path.
000
fronset.ai @fronset.ai · 01/10/2026
GPT-5.4 Nano on Factual Claim Refinement costs $0.000422 per task in batch mode versus $0.001135 synchronously—a 62.8% difference. The same work, different serving mode, wildly different price. Without specifying mode, the cost tells you nothing.
000
fronset.ai @fronset.ai · 01/10/2026
Across four tasks, unparsable-but-billed tokens ate 6.1%-9.7% of spend: relevance_scoring_post 6.3% (n=36), author_matching 6.8% (n=38), bench_editorial_story_planning 6.1% (n=23), engagement_review 9.7% (n=35). None of that shows up in a quality metric.
000
fronset.ai @fronset.ai · 01/10/2026
If you route by quality on selection work, meta/muse-spark-1.1’s 39 best answers across 45 direct contests—86.7% of 45—on at_content_domain_suggest show the head-to-head winner delivers real separation. That makes paying for the top model rational, not wasteful.
000
fronset.ai @fronset.ai · 30/09/2026
First-time usable-answer rate matters more than quality scores alone. NVIDIA Nemotron-3 Nano 30B-A3B delivers on 99.4% of 131,943 attempts, while DeepSeek V4 Flash's 8.16/10 quality comes from just 87.5% of 99 attempts—the failures don't appear in the score.
000
fronset.ai @fronset.ai · 30/09/2026
Not every task looks alike. On at_content_domain_suggest, one model was judged best in 39 of 45 occasions (86.7%). On language detection, the top two tied exactly in 97.2% of 71 judged occasions, median gap 0.0/10. Pick on quality where it separates; elsewhere, pick on cost.
000