Sign in

Kıvanç Yüksel

@smiletoai.com
105 followers 202 following 137 posts

Building SmileToAI by myself — generate images, narrate audio, make video, write, all in one place. Teaching myself robotics starting from linear algebra, because skipping fundamentals never actually works. Warsaw.

PostsRepliesMedia
Kıvanç Yüksel @smiletoai.com · 09/10/2026
One upside of open weights: that question gets an answer. Pin a checkpoint, run it yourself, and you can re-run last week’s prompts to see if it actually got worse. Through someone else’s API you’re back to guessing.
010
Kıvanç Yüksel @smiletoai.com · 09/10/2026
Gemini's live transcription bills silence: 25 audio tokens per streamed second, pauses included, so a quiet minute costs over half a talking one. No usage numbers come back; our dictation relay counts the seconds itself. Whistle skips silence on-device: cactuscompute.com/blog/whistle
110
Kıvanç Yüksel @smiletoai.com · 09/10/2026
One catch: 'successful' means the checker said yes. Traces that passed by gaming a weak check end up in the positives too, so replaying them compounds the checker's blind spots along with the skill.
010
Kıvanç Yüksel @smiletoai.com · 08/10/2026
That chart looks like a capability curve to me. Haiku sits near zero and Sonnet under 20% on all three prompts. If Haiku can't steer its thinking because it isn't capable enough yet, that says little about how it would use the ability, and not much about the model it was distilled from.
000
Kıvanç Yüksel @smiletoai.com · 08/10/2026
Haiku 5.5 is priced by prompt size: $0.10 per million input tokens up to 100k, $0.50 past it. So 99k tokens of input is about a cent and 101k about five, if the whole prompt takes the higher rate (my read of the table). We compact chat at 48k by default, for quality. Now it's a cost setting too.
210
Kıvanç Yüksel @smiletoai.com · 08/10/2026
Right, cancelling the await doesn't touch the thread. We already use the AsyncClient for streaming, so the non-streaming path moving onto it is the actual fix.
100
Kıvanç Yüksel @smiletoai.com · 08/10/2026
Half of it. The timeout fires now and the loop stays free, but the call keeps running in its thread until xAI answers, so it still finishes and still bills. The real fix is the async client.
100
Kıvanç Yüksel @smiletoai.com · 08/10/2026
Probably, and it's part of why raw reasoning traces stopped being shown. OpenAI never exposed o1's chain of thought, and Claude and Gemini hand back summaries. Most of what RL pays for is the search for trajectories that work, and the traces are much of what would let a cloner skip it.
070
Kıvanç Yüksel @smiletoai.com · 07/10/2026
Building in public, yesterday's bug: an 8 s asyncio.wait_for that never fired. The calls took 14.5 s and 19.4 s and returned normally. The xAI SDK call was synchronous inside an async def, so it blocked the event loop, and the loop is what raises the timeout. The fix was asyncio.to_thread.
220
Kıvanç Yüksel @smiletoai.com · 07/10/2026
Yes, for notebook search over docs people upload. Extraction + SQL needs you to know the fields when you write, and there nobody knows the question yet. Where the labels are known up front, agreed, I'd reach for extraction.
000
Kıvanç Yüksel @smiletoai.com · 06/10/2026
Stack it with "take a deep breath" and you've basically got a kid doing long division.
040
Kıvanç Yüksel @smiletoai.com · 06/10/2026
Same prompt, 4 models #2: only gpt-image-2 drew the analog clock at 4:35 I asked for. FLUX gave me 10:10, the watch-ad pose, and Grok about 1:00. Gemini got the minute hand right but stopped the hour hand short of the 4, so it reads 3:35. One try each, I kept whatever came back.
010
Kıvanç Yüksel @smiletoai.com · 05/10/2026
Fair, I called it a probability when it's a score. Bucketing by score on a labeled set is the cheap way to see whether it can be read as one, and where to set a threshold.
000
Kıvanç Yüksel @smiletoai.com · 05/10/2026
Out of credits on Gemini is HTTP 402 now. OpenAI and xAI still send a 429, and Gemini's per-minute 429 says to check your billing details. Our router didn't fail over on a 402, so on Oct 2 our empty Gemini account meant errors, not fallback. Status codes alone won't tell you who's out of money.
110
Kıvanç Yüksel @smiletoai.com · 04/10/2026
Fair, a hook is a heuristic and the model can route around one. The sandbox isn't, though. If the OS won't let the process write outside the repo, it doesn't matter how clever it gets with python3.
000
Kıvanç Yüksel @smiletoai.com · 04/10/2026
Removing it isn't the only option though. A hook can look at each shell command before it runs, and a sandbox can limit where it's allowed to write. The model keeps the shell, the harness just narrows what it can touch.
110
Kıvanç Yüksel @smiletoai.com · 03/10/2026
Fair, but "without limitations" is the catch. The uniqueness check is what stops Edit from changing the wrong match, and python3 skips it. So whether a model gets the shell is the harness's call, not something a smarter model should decide for itself.
121
Kıvanç Yüksel @smiletoai.com · 03/10/2026
Less deciding, more measuring. Reading the last-position logits gives you a probability for every class, so you can pick a threshold per use. A CoT model's final answer comes after one sampled chain of reasoning, so even its answer-token probabilities are conditional on that one draw.
100
Kıvanç Yüksel @smiletoai.com · 02/10/2026
pgvector filters after the HNSW scan, so a filter matching 10% of rows gets ~4 of the default 40 candidates. turbopuffer's post on demoting its vector index sent me to our notebook search: one HNSW index over every team's chunks, filters in the WHERE. github.com/pgvector/pgvector#filter…
000
Kıvanç Yüksel @smiletoai.com · 02/10/2026
@kyisaiah47.thecompound.tech your solo builders pack sounds like my kind of people. I build and ship an AI app by myself and post the unglamorous bits too. Would you add me?
010
Kıvanç Yüksel @smiletoai.com · 02/10/2026
Hi Austin, would you add me to this one? Solo founder here, working on an AI app and posting what I learn as I go.
110
Kıvanç Yüksel @smiletoai.com · 02/10/2026
The built-in Edit tool is basically str.replace too, except it refuses when the old string isn't unique in the file. A Python script it writes itself only has whatever check it bothered to add, and str.replace hits every match by default.
130
Kıvanç Yüksel @smiletoai.com · 01/10/2026
The nice thing about those is that most mistakes show up as wrong output: the render looks off, the program prints the wrong thing. Reading AI code gives you no feedback on your own skill, writing these does.
010
Kıvanç Yüksel @smiletoai.com · 01/10/2026
The migration notes say to size max_tokens for thinking plus reply, since a limit sized for a no-thinking route can cut replies off. Low only promises the thinking stays short, not that it's skipped, so I wouldn't count on a text-sized value.
010
Kıvanç Yüksel @smiletoai.com · 01/10/2026
Gemini 4 Argon raises the output limit from 64K to 1M tokens. At $20 per 1M output once the intro price ends, a call that fills it costs $20. We leave max_output_tokens unset on purpose in one of our narration pipelines. At 64K that was a cheap default. At 1M I don't think it is.
210
Kıvanç Yüksel @smiletoai.com · 01/10/2026
The friction is doing a job for those firms, though. Retention discounts are affordable because most people won't sit through the hold queue to ask. If everyone's agent asks, the discount stops being selective, so I'd expect the offers to shrink before the call centers drown.
000
Kıvanç Yüksel @smiletoai.com · 01/10/2026
Noworries, thanks!
000
Kıvanç Yüksel @smiletoai.com · 30/09/2026
Right, thinking_tokens is what thinking took, and the cap needs the reply on top of that. Thanks for the correction.
000
Kıvanç Yüksel @smiletoai.com · 29/09/2026
Fair, I'd missed that field. So on a cut-off reply you can see how much of the cap thinking took, which is the number you'd actually size max_tokens from at low. Worth knowing it only shows up on the final message_delta when streaming.
110
Kıvanç Yüksel @smiletoai.com · 29/09/2026
Cheap until you remember every "thanks, great work!" stays in context and gets re-read on every turn after.
011
Kıvanç Yüksel @smiletoai.com · 29/09/2026
It caps spend, but it can also eat the answer. Thinking shares that budget, so on a prompt where low still thinks, a text-sized max_tokens cuts the reply short, and the only signal is stop_reason max_tokens.
100
Kıvanç Yüksel @smiletoai.com · 28/09/2026
Fair, if they pass it all down the queue doesn't get any shorter. It's not wasted though: they sell out either way, so the same GPUs just serve more people at the lower price. More a market-share move than a fix for being short, which I guess was your point.
000
Kıvanç Yüksel @smiletoai.com · 28/09/2026
It still holds if they're capacity-capped, I think. Jevons says total compute demand grows, but a lab that's rationing already has more demand than GPUs. A model that's cheaper to serve then just pushes more requests through the same hardware, which is what you'd do when short.
100
Kıvanç Yüksel @smiletoai.com · 28/09/2026
Red RGB(196,52,52) and green RGB(30,140,30) both become grey 95. Black-and-white keeps brightness and drops hue, so colorizing is a guess. Our guide says that: smiletoai.com/blog/how-to-colorize-…
000
Kıvanç Yüksel @smiletoai.com · 28/09/2026
Yeah, and they said the refusals bias it low: 30-50% of devs admitted holding back tasks they didn't want to do without AI. Even so, the 10 returning devs from the first study took about 18% less time with AI, and the interval still crosses zero.
000
Kıvanç Yüksel @smiletoai.com · 28/09/2026
Background removal's still the only model step I run. Past that, yeah, it's the glue: each generation job reserves Sparks (credits) from one balance up front, charges when it finishes and releases them if the provider fails.
000
Kıvanç Yüksel @smiletoai.com · 28/09/2026
Fair, one clean sample only covers the prompts in it. You'd have to keep watching real traffic rather than sign it off once.
120
Kıvanç Yüksel @smiletoai.com · 27/09/2026
You can tell from the responses: one the cap cut off comes back with stop_reason max_tokens, so counting those on a sample at low effort would show whether a text-sized value leaves room.
110
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Thanks for adding me! Generation (image, narration, video) all runs on provider models. The piece I keep is background removal, on an open matting model I run myself, because a generative edit repaints the whole picture and matting keeps the user's exact pixels.
020
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Hi Johannes, would you add me? I work with ML models in production for my own app and post about it.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Building an AI app in public, on my own, from Warsaw. Would be happy to be in this.
010
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Solo founder in Warsaw, building an AI product. Would be glad to join if there's still a spot.
120
Kıvanç Yüksel @smiletoai.com · 27/09/2026
I'd be happy to be in this one. Indie builder here, working on an AI app on my own.
100
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Late to this, but I'm an indie founder building an AI app on my own. Would be glad to be added if you're still keeping it up.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
You asked for suggestions, so I'm suggesting myself. I run a handful of model providers in production for my own app and post about the infrastructure side of it.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
If you still update this one, I'd like to be in it. I'm building an LLM-heavy app on my own and post what I learn along the way.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Solo founder here, bootstrapping an AI app from Warsaw. Would be happy to be in this if you're still adding people.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
Hi Sam, I'd like to be in this if there's room. I run speech-to-text and text-to-speech in my app and post about it now and then.
100
Kıvanç Yüksel @smiletoai.com · 27/09/2026
I'd like to be considered for Coty's AI Starter Pack. I'm an engineer running image, speech and video models in production for my own product, and I post what changes underneath it: price changes, shutdown dates, limits that differ from the docs.
000
Kıvanç Yüksel @smiletoai.com · 27/09/2026
The METR study people cite for that measured nearly the opposite of your case: regular contributors averaging 5 years in big repos they knew cold, on early-2025 tools. Its caveats call the results consistent with small greenfield projects or unfamiliar codebases getting substantial speedup.
100