Kıvanç Yüksel @smiletoai.com · 09/10/2026One upside of open weights: that question gets an answer. Pin a checkpoint, run it yourself, and you can re-run last week’s prompts to see if it actually got worse. Through someone else’s API you’re back to guessing. 010
Kıvanç Yüksel @smiletoai.com · 09/10/2026Gemini's live transcription bills silence: 25 audio tokens per streamed second, pauses included, so a quiet minute costs over half a talking one. No usage numbers come back; our dictation relay counts the seconds itself. Whistle skips silence on-device: cactuscompute.com/blog/whistle 110
Kıvanç Yüksel @smiletoai.com · 09/10/2026One catch: 'successful' means the checker said yes. Traces that passed by gaming a weak check end up in the positives too, so replaying them compounds the checker's blind spots along with the skill. 010
Kıvanç Yüksel @smiletoai.com · 08/10/2026That chart looks like a capability curve to me. Haiku sits near zero and Sonnet under 20% on all three prompts. If Haiku can't steer its thinking because it isn't capable enough yet, that says little about how it would use the ability, and not much about the model it was distilled from. 000
Kıvanç Yüksel @smiletoai.com · 08/10/2026Haiku 5.5 is priced by prompt size: $0.10 per million input tokens up to 100k, $0.50 past it. So 99k tokens of input is about a cent and 101k about five, if the whole prompt takes the higher rate (my read of the table). We compact chat at 48k by default, for quality. Now it's a cost setting too. 210
Kıvanç Yüksel @smiletoai.com · 08/10/2026Right, cancelling the await doesn't touch the thread. We already use the AsyncClient for streaming, so the non-streaming path moving onto it is the actual fix. 100
Kıvanç Yüksel @smiletoai.com · 08/10/2026Half of it. The timeout fires now and the loop stays free, but the call keeps running in its thread until xAI answers, so it still finishes and still bills. The real fix is the async client. 100
Kıvanç Yüksel @smiletoai.com · 08/10/2026Probably, and it's part of why raw reasoning traces stopped being shown. OpenAI never exposed o1's chain of thought, and Claude and Gemini hand back summaries. Most of what RL pays for is the search for trajectories that work, and the traces are much of what would let a cloner skip it. 070
Kıvanç Yüksel @smiletoai.com · 07/10/2026Building in public, yesterday's bug: an 8 s asyncio.wait_for that never fired. The calls took 14.5 s and 19.4 s and returned normally. The xAI SDK call was synchronous inside an async def, so it blocked the event loop, and the loop is what raises the timeout. The fix was asyncio.to_thread. 220
Kıvanç Yüksel @smiletoai.com · 07/10/2026Yes, for notebook search over docs people upload. Extraction + SQL needs you to know the fields when you write, and there nobody knows the question yet. Where the labels are known up front, agreed, I'd reach for extraction. 000
Kıvanç Yüksel @smiletoai.com · 06/10/2026Stack it with "take a deep breath" and you've basically got a kid doing long division. 040
Kıvanç Yüksel @smiletoai.com · 06/10/2026Same prompt, 4 models #2: only gpt-image-2 drew the analog clock at 4:35 I asked for. FLUX gave me 10:10, the watch-ad pose, and Grok about 1:00. Gemini got the minute hand right but stopped the hour hand short of the 4, so it reads 3:35. One try each, I kept whatever came back. 010
Kıvanç Yüksel @smiletoai.com · 05/10/2026Fair, I called it a probability when it's a score. Bucketing by score on a labeled set is the cheap way to see whether it can be read as one, and where to set a threshold. 000
Kıvanç Yüksel @smiletoai.com · 05/10/2026Out of credits on Gemini is HTTP 402 now. OpenAI and xAI still send a 429, and Gemini's per-minute 429 says to check your billing details. Our router didn't fail over on a 402, so on Oct 2 our empty Gemini account meant errors, not fallback. Status codes alone won't tell you who's out of money. 110
Kıvanç Yüksel @smiletoai.com · 04/10/2026Fair, a hook is a heuristic and the model can route around one. The sandbox isn't, though. If the OS won't let the process write outside the repo, it doesn't matter how clever it gets with python3. 000
Kıvanç Yüksel @smiletoai.com · 04/10/2026Removing it isn't the only option though. A hook can look at each shell command before it runs, and a sandbox can limit where it's allowed to write. The model keeps the shell, the harness just narrows what it can touch. 110
Kıvanç Yüksel @smiletoai.com · 03/10/2026Fair, but "without limitations" is the catch. The uniqueness check is what stops Edit from changing the wrong match, and python3 skips it. So whether a model gets the shell is the harness's call, not something a smarter model should decide for itself. 121
Kıvanç Yüksel @smiletoai.com · 03/10/2026Less deciding, more measuring. Reading the last-position logits gives you a probability for every class, so you can pick a threshold per use. A CoT model's final answer comes after one sampled chain of reasoning, so even its answer-token probabilities are conditional on that one draw. 100
Kıvanç Yüksel @smiletoai.com · 02/10/2026pgvector filters after the HNSW scan, so a filter matching 10% of rows gets ~4 of the default 40 candidates. turbopuffer's post on demoting its vector index sent me to our notebook search: one HNSW index over every team's chunks, filters in the WHERE. github.com/pgvector/pgvector#filter… 000
Kıvanç Yüksel @smiletoai.com · 02/10/2026@kyisaiah47.thecompound.tech your solo builders pack sounds like my kind of people. I build and ship an AI app by myself and post the unglamorous bits too. Would you add me? 010
Kıvanç Yüksel @smiletoai.com · 02/10/2026Hi Austin, would you add me to this one? Solo founder here, working on an AI app and posting what I learn as I go. 110
Kıvanç Yüksel @smiletoai.com · 02/10/2026The built-in Edit tool is basically str.replace too, except it refuses when the old string isn't unique in the file. A Python script it writes itself only has whatever check it bothered to add, and str.replace hits every match by default. 130
Kıvanç Yüksel @smiletoai.com · 01/10/2026The nice thing about those is that most mistakes show up as wrong output: the render looks off, the program prints the wrong thing. Reading AI code gives you no feedback on your own skill, writing these does. 010
Kıvanç Yüksel @smiletoai.com · 01/10/2026The migration notes say to size max_tokens for thinking plus reply, since a limit sized for a no-thinking route can cut replies off. Low only promises the thinking stays short, not that it's skipped, so I wouldn't count on a text-sized value. 010
Kıvanç Yüksel @smiletoai.com · 01/10/2026Gemini 4 Argon raises the output limit from 64K to 1M tokens. At $20 per 1M output once the intro price ends, a call that fills it costs $20. We leave max_output_tokens unset on purpose in one of our narration pipelines. At 64K that was a cheap default. At 1M I don't think it is. 210
Kıvanç Yüksel @smiletoai.com · 01/10/2026The friction is doing a job for those firms, though. Retention discounts are affordable because most people won't sit through the hold queue to ask. If everyone's agent asks, the discount stops being selective, so I'd expect the offers to shrink before the call centers drown. 000
Kıvanç Yüksel @smiletoai.com · 30/09/2026Right, thinking_tokens is what thinking took, and the cap needs the reply on top of that. Thanks for the correction. 000
Kıvanç Yüksel @smiletoai.com · 29/09/2026Fair, I'd missed that field. So on a cut-off reply you can see how much of the cap thinking took, which is the number you'd actually size max_tokens from at low. Worth knowing it only shows up on the final message_delta when streaming. 110
Kıvanç Yüksel @smiletoai.com · 29/09/2026Cheap until you remember every "thanks, great work!" stays in context and gets re-read on every turn after. 011
Kıvanç Yüksel @smiletoai.com · 29/09/2026It caps spend, but it can also eat the answer. Thinking shares that budget, so on a prompt where low still thinks, a text-sized max_tokens cuts the reply short, and the only signal is stop_reason max_tokens. 100
Kıvanç Yüksel @smiletoai.com · 28/09/2026Fair, if they pass it all down the queue doesn't get any shorter. It's not wasted though: they sell out either way, so the same GPUs just serve more people at the lower price. More a market-share move than a fix for being short, which I guess was your point. 000
Kıvanç Yüksel @smiletoai.com · 28/09/2026It still holds if they're capacity-capped, I think. Jevons says total compute demand grows, but a lab that's rationing already has more demand than GPUs. A model that's cheaper to serve then just pushes more requests through the same hardware, which is what you'd do when short. 100
Kıvanç Yüksel @smiletoai.com · 28/09/2026Red RGB(196,52,52) and green RGB(30,140,30) both become grey 95. Black-and-white keeps brightness and drops hue, so colorizing is a guess. Our guide says that: smiletoai.com/blog/how-to-colorize-… 000
Kıvanç Yüksel @smiletoai.com · 28/09/2026Yeah, and they said the refusals bias it low: 30-50% of devs admitted holding back tasks they didn't want to do without AI. Even so, the 10 returning devs from the first study took about 18% less time with AI, and the interval still crosses zero. 000
Kıvanç Yüksel @smiletoai.com · 28/09/2026Background removal's still the only model step I run. Past that, yeah, it's the glue: each generation job reserves Sparks (credits) from one balance up front, charges when it finishes and releases them if the provider fails. 000
Kıvanç Yüksel @smiletoai.com · 28/09/2026Fair, one clean sample only covers the prompts in it. You'd have to keep watching real traffic rather than sign it off once. 120
Kıvanç Yüksel @smiletoai.com · 27/09/2026You can tell from the responses: one the cap cut off comes back with stop_reason max_tokens, so counting those on a sample at low effort would show whether a text-sized value leaves room. 110
Kıvanç Yüksel @smiletoai.com · 27/09/2026Thanks for adding me! Generation (image, narration, video) all runs on provider models. The piece I keep is background removal, on an open matting model I run myself, because a generative edit repaints the whole picture and matting keeps the user's exact pixels. 020
Kıvanç Yüksel @smiletoai.com · 27/09/2026Hi Johannes, would you add me? I work with ML models in production for my own app and post about it. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026Building an AI app in public, on my own, from Warsaw. Would be happy to be in this. 010
Kıvanç Yüksel @smiletoai.com · 27/09/2026Solo founder in Warsaw, building an AI product. Would be glad to join if there's still a spot. 120
Kıvanç Yüksel @smiletoai.com · 27/09/2026I'd be happy to be in this one. Indie builder here, working on an AI app on my own. 100
Kıvanç Yüksel @smiletoai.com · 27/09/2026Late to this, but I'm an indie founder building an AI app on my own. Would be glad to be added if you're still keeping it up. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026You asked for suggestions, so I'm suggesting myself. I run a handful of model providers in production for my own app and post about the infrastructure side of it. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026If you still update this one, I'd like to be in it. I'm building an LLM-heavy app on my own and post what I learn along the way. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026Solo founder here, bootstrapping an AI app from Warsaw. Would be happy to be in this if you're still adding people. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026Hi Sam, I'd like to be in this if there's room. I run speech-to-text and text-to-speech in my app and post about it now and then. 100
Kıvanç Yüksel @smiletoai.com · 27/09/2026I'd like to be considered for Coty's AI Starter Pack. I'm an engineer running image, speech and video models in production for my own product, and I post what changes underneath it: price changes, shutdown dates, limits that differ from the docs. 000
Kıvanç Yüksel @smiletoai.com · 27/09/2026The METR study people cite for that measured nearly the opposite of your case: regular contributors averaging 5 years in big repos they knew cold, on early-2025 tools. Its caveats call the results consistent with small greenfield projects or unfamiliar codebases getting substantial speedup. 100