Sign in

Damien Tournoud

@damientournoud.bsky.social
64 followers 18 following 89 posts
PostsRepliesMedia
Damien Tournoud @damientournoud.bsky.social · 30/04/2026
It would have helped with this, I'm sure. You want your model to keep a cool head. bsky.app/profile/dami...
000
Damien Tournoud @damientournoud.bsky.social · 30/04/2026
Makes total sense
210
Reposted by Damien Tournoud
Why @why.bsky.world · 30/04/2026
Gemma 4 runs best with liquid cooling
6663378
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
6/ The gemma4 case is stranger. `<|channel>` tags would have to be fabricated by a client that already speaks Gemma's internal vocabulary. No reasonable client does that. The renderer still guards against it, apparently because the upstream Hugging Face chat template did.
000
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
5/ The strip only matters when a *client* does something weird: stuffs internal tokens into `content` where they don't belong. Open-WebUI is one: it inlines past thinking into content as `<think>...</think>`, which is exactly the shape `filterThinkTags` hunts for.
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
4/ Here's the puzzle. Both `<think>...</think>` and `<|channel>...<channel|>` are the model's *internal* token format. ollama's API never returns them in `content`. `api.Message` has separate `content` + `thinking` JSON fields, parsed apart before the response is serialized.
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
3/ Place two: stripThinking in the gemma4 renderer walks past assistant content and removes `<|channel>...<channel|>` blocks before re-rendering. Comment says it mirrors Hugging Face's `strip_thinking` macro.
model/renderers/gemma4.go:
// stripThinking removes <|channel>...<channel|> thinking blocks from content,
// matching the HF chat template's strip_thinking macro.
func stripThinking(text string) string {
	var result strings.Builder
	for {
		start := strings.Index(text, "<|channel>")
		if start == -1 {
			result.WriteString(text)
			break
		}
		result.WriteString(text[:start])
		end := strings.Index(text[start:], "<channel|>")
		if end == -1 {
			break
		}
		text = text[start+end+len("<channel|>"):]
	}
	return strings.TrimSpace(result.String())
}
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
2/ Place one: filterThinkTags in `server/routes.go` walks past assistant messages, runs each through a `<think>`/`</think>` parser, and rewrites `Content` to just the post-thinking remainder. Only fires for qwen3 family and deepseek-r1.
server/routes.go:
func filterThinkTags(msgs []api.Message, m *Model) []api.Message {
	if m.Config.ModelFamily == "qwen3" || model.ParseName(m.Name).Model == "deepseek-r1" {
		finalUserIndex := -1
		for i, msg := range msgs {
			if msg.Role == "user" {
				finalUserIndex = i
			}
		}

		for i, msg := range msgs {
			if msg.Role == "assistant" && i < finalUserIndex {
				// TODO(drifkin): this is from before we added proper thinking support.
				// However, even if thinking is not enabled (and therefore we shouldn't
				// change the user output), we should probably perform this filtering
				// for all thinking models (not just qwen3 & deepseek-r1) since it tends
				// to save tokens and improve quality.
				thinkingState := &thinking.Parser{
					OpeningTag: "<think>",
					ClosingTag: "</think>",
				}
				_, content := thinkingState.AddContent(msg.Content)
				msgs[i].Content = content
			}
		}
	}
	return msgs
}
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
Follow-up to this morning's rabbit hole. Ollama strips thinking-trace markers from *input* messages in two different places, but the markers it strips never appear in its own API output to begin with. That's very puzzling. 🧵
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
17/ Lingering puzzle: same Open-WebUI setup with the dense gemma4:31b never reproduced. Only the Mixture-of-Experts 26B did. My hypothesis (n=1): MoE models are unusually sensitive to format cues in context, while a dense model's format habits are strong enough to ignore the priming.
010
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
16/ Of course, a blanket strip doesn't work for models that specifically need some reasoning reflected back (Claude's extended thinking for tool-use continuity, Gemini's thought signatures, etc.). Those will need a different approach, where model-specific metadata is stored alongside the content.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
15/ Workaround until #23339 ships: an Open-WebUI Filter Function with an `inlet` hook that walks past assistant messages and strips thinking tags before they're sent back to the model. Someone already wrote one.
gist.github.com
Open WebUI filter to prevent thinking re-injection
Open WebUI filter to prevent thinking re-injection - filter.py
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
14/ The odd thing is that the convention is universal across reasoning models: thinking is per-turn and never persisted in history. Qwen3 docs say strip it. DeepSeek-R1 says strip it. o1 doesn't even return it. Open-WebUI is the side breaking the convention.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
13/ Open issue: open-webui/open-webui#23339. Reinjecting thinking causes models to imitate and amplify the tags. A model-level toggle is planned, default off. Not yet shipped.
github.com
feat: Add model-level toggle to disable reinjecting reasoning/thinking into prompts (prevents <think> tag imitation and rendering/parsing failures) · Issue #23339 · open-webui/open-webui
Check Existing Issues I have searched for all existing open AND closed issues and discussions for similar requests. I have found none that is comparable to my request. Verify Feature Scope I have r...
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
12/ Where did `<think>...</think>` come from in past assistant content? Open-WebUI. Its chat format has no separate thinking field, so it inlines thinking into content using DeepSeek-R1's `<think>{thinking}</think>{content}` convention and ships that on every turn.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
11/ Why was the model emitting `<think>...</think>` instead of its native tags? The chat history sent on every turn contained `<think>...</think>` blocks in past assistant messages. The model was imitating what it saw.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
10/ So here is what actually happened: the model produced a clean answer, but it closed its thinking with `</think>`, not its native `<channel|>` token. ollama doesn't recognize `</think>`. Everything after that "close" stayed in the thinking buffer. Content empty. Bug reproduced.
120
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
9/ (Of course I could have used Charles or another off-the-shelf debugging proxy. In this case it was just easier to ask Claude to write one instead of going through the trouble of finding and configuring one.)
120
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
8/ So I got Claude Code to write a tiny logging proxy (~150 lines of Python) between Open-WebUI and ollama, capturing every request body and the full streamed response. Pointed Open-WebUI at it. The reason became obvious immediately.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
7/ I stress-tested 150+ runs across various prompts and long contexts. Zero reproductions on anything I could synthesize. I could trigger it from the application but my synthetic harness couldn't. Setup difference mattered.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
6/ I had Claude Code cross-reference ollama's gemma4 implementation against the reference one in Hugging Face transformers. They match. The runner code seems correct.
github.com
transformers/src/transformers/models/gemma4/modular_gemma4.py at main · huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training. - huggingface/transformers
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
5/ I disproved that by looking at how likely each token actually was across the run. The suspect tokens never came anywhere close to being selected. The model wasn't being clipped early.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
4/ My first wrong guess: the model was hitting an end-of-output marker mid-thinking by accident. gemma4's stop list does include `<|tool_response>` (the *opening* tag of a tool response) as a guardrail, so it was at least plausible.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
3/ gemma4 wraps its thinking in custom tokens: `<|channel>thought\n...<channel|>`. ollama watches for that specific close tag and splits thinking from content there. If anything else shows up in its place, everything spills into the thinking bucket and content stays empty.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
2/ Symptom: a multi-paragraph "thinking" block that quietly contained the actual answer somewhere inside, then an empty reply. ollama captured the entire output as thinking and never noticed it was supposed to stop.
110
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
Rabbit holes are always fun. Spent a few hours debugging why gemma4:26b (MoE) on ollama via Open-WebUI kept returning replies that looked empty — long internal thinking, then nothing. Only the MoE reproduced but the dense 31b never did. A debugging 🧵
111
Damien Tournoud @damientournoud.bsky.social · 18/03/2026
Woo, huge milestone. Great job @dholms.at!
050
Damien Tournoud @damientournoud.bsky.social · 16/03/2026
In the best of times, getting Claude Code to follow instructions is more art than science, but with Opus 4.6 (1M context), it is just wild how bad it is when you reach even 50% of the context window.
011
Damien Tournoud @damientournoud.bsky.social · 11/03/2026
What could possibly be happening in that LLM to get this garbage as an output?
Google AI overview for the search "strait of hormuz gps interference"

The model output is all garbled, with commas all over the place, like:

"These COMMA electronic warfare attacks create COMMA dangerous navigation COMMA conditions COMMA leading to COMMA ship collisions COMMA drifting COMMA and COMMA skyrocketing insurance COMMA premiums for COMMA carriers"Continued garbled outputContinue garbled output
000
Damien Tournoud @damientournoud.bsky.social · 06/03/2026
Amazing
010
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
Yup. It can be very quick to prematurely declare Mission Accomplished. I have also seen it randomly decide not to do something and declare victory anyway.
Picture of the George W. Bush's 2003 "Mission Accomplished" speech
000
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
Claude Code was a pretty solid force multiplier on this. But I find that a lot of the intelligence is between the keyboard and the chair. It helps a lot to know exactly what you want...
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
9/ It's a single Go binary: gmailfs -mountpoint /mnt/gmail Handles OAuth2 on first run, then you're browsing. Go with go-fuse for the filesystem layer and PebbleDB for caching.
github.com
GitHub - damz/gmailfs: Mount your Gmail as a local directory tree organized by label and date, with emails as .eml files.
Mount your Gmail as a local directory tree organized by label and date, with emails as .eml files. - damz/gmailfs
000
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
8/ Oh, and surprisingly Gmail's API has no "All Mail" label. In the UI, it's just "everything not in Trash or Spam." So in gmailfs we synthesizes one by querying without a label filter.
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
7/ FOPEN_CACHE_DIR on OPENDIR tells the kernel to cache directory listings. But OPENDIR/RELEASEDIR are still two extra roundtrips per listing. go-fuse didn't support opting out, so I sent a patch upstream to add FUSE_NO_OPENDIR_SUPPORT — READDIR without a file handle, no unnecessary roundtrips.
review.gerrithub.io
fs: support READDIR/READDIRPLUS with no file handle (FUSE_NO_OPENDIR_SUPPORT)
Gerrit Code Review
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
6/ The kernel FUSE layer caches too. Infinite attribute/entry/negative timeouts, OPENDIR/RELEASEDIR opt-out so the kernel caches READDIR across opens. When history sync detects changes, targeted NotifyEntry calls invalidate only the affected paths.
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
5/ Gmail's API is rate limited, so gmailfs caches aggressively in PebbleDB. Background history sync detects changes and invalidates only what's needed — usually just a single day. The year and month structure almost never needs to be recomputed.
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
4/ Figuring out which years have messages without listing everything: scan backwards from the newest message. Record its year, jump to the start of that year, find the next oldest. Same for months and days. Empty periods cost zero API calls.
gmail.go:
func (g *GmailClient) PopulatedYears(ctx context.Context, labelID string, cache *Cache) ([]int, error) {
	var years []int
	cursor := time.Now().AddDate(1, 0, 0).Unix()
	for {
		s, found, err := g.newestMessageInRange(ctx, labelID, 0, cursor, cache)
		if err != nil {
			return nil, err
		}

		if !found {
			break
		}

		y := time.UnixMilli(s.InternalDate).In(time.Local).Year()
		years = append(years, y)
		cursor = time.Date(y, 1, 1, 0, 0, 0, 0, time.Local).Unix()
	}

	slices.Reverse(years)
	return years, nil
}
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
3/ Every label becomes a folder, every email a .eml file, organized by year/month/day. Only populated directories are shown — no scrolling past empty months.
/mnt/gmail/:
All Mail/
  2024/
    01/
      15/
        2024-01-15T143022-Meeting-notes.eml
        2024-01-15T091205-Your-order-has-shipped.eml
Inbox/
  2026/
    02/
      27/
        2026-02-27T083412-Re-Project-update.eml
        2026-02-27T071055-Weekly-digest.eml
Sent/
  2026/
    02/
      26/
        2026-02-26T154830-Re-Lunch-tomorrow.eml
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
2/ By making email look like files, every tool that reads files works for free. Claude Code browses with ls, greps for keywords, reads messages with cat. No MCP server needed, it just uses its built-in tools.
110
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
I wanted to ask Claude Code questions about my email. "What did I discuss with X last week?" "Find me that tracking number from yesterday." Turns out, if you mount Gmail as a FUSE filesystem, Claude Code already knows how to browse and search files. So I built gmailfs. 🧵
github.com
GitHub - damz/gmailfs: Mount your Gmail as a local directory tree organized by label and date, with emails as .eml files.
Mount your Gmail as a local directory tree organized by label and date, with emails as .eml files. - damz/gmailfs
221
Damien Tournoud @damientournoud.bsky.social · 16/02/2026
The Verge is trying to explain quoted-printable. The lone equal signs are soft line wraps at 76 characters. But why do they seem to eat the next character? Maybe they tried to hand-decode the quoted-printable encoding (maybe with a regexp matching "=.."?), before an HTML-to-text conversion step?
000
Damien Tournoud @damientournoud.bsky.social · 21/01/2026
That's a lot of data movement...
010
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
Long story short, update to Golang 1.25.6 or 1.24.12.
000
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
In my testing, with just 2 reqs/s (concurrency of 10) you can easily get the server to allocate 8 GB+ and consume 15 cores of CPU. This is even worse when the server supports compressed request bodies (but my sense is that it is uncommon), for example with the Decompress middleware in Echo.
100
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
This results in the allocation of a map with 2.1 million entries, and the corresponding 2.1 million short key strings. In total, 416 MB of allocation.
Parsing a 10 MB body results in 416 MB of allocation over 2.1 million allocations:
$ go test -v -bench=. -benchmem .
BenchmarkMemory-20             1        1126808685 ns/op
        437066424 B/op   2142306 allocs/op
100
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
The issue is a classic memory amplification issue: while Request.ParseForm refuses by default to parse request bodies bigger than 10 MB, even that can result in significant memory usage for specially crafted inputs. In this case, the input is a URL-encoded form data with small keys and no values.
Example application/x-www-form-urlencoded body:
a&b&c&d&e&f&g&h&i&j&k&l&m&n&o&p&q&r&s&t&u&v&w&x&y&z&A&B&C&D&E&F&G&H&I&J&K&L
&M&N&O&P&Q&R&S&T&U&V&W&X&Y&Z&ab&bb&cb&db&eb&fb&gb&hb&ib&jb&kb&lb&mb&nb&ob&p
b&qb&rb&sb&tb&ub&vb&wb&xb&yb&zb&Ab&Bb&Cb&Db&Eb&Fb&Gb&Hb&Ib&Jb&Kb&Lb&Mb&Nb&O
b&Pb&Qb&Rb&Sb&Tb&Ub&Vb&Wb&Xb&Yb&Zb&ac&bc&cc&dc&ec&fc&gc&hc&ic&jc&kc&lc&mc&n
c&oc&pc&qc&rc&sc&tc&uc&vc&wc&xc&yc&zc&Ac&Bc&Cc&Dc&Ec&Fc&Gc&Hc&Ic&Jc&Kc&Lc&M
c&Nc&Oc&Pc&Qc&Rc&Sc&Tc&Uc&Vc&Wc&Xc&Yc&Zc&ad&bd&cd&dd&ed&fd&gd&hd&id&jd&kd&l
d&md&nd&od&pd&qd&rd&sd&td&ud&vd&wd&xd&yd&zd&Ad&Bd&Cd&Dd&Ed&Fd&Gd&Hd&Id&Jd&K
d&Ld&Md&Nd&Od&Pd&Qd&Rd&Sd&Td&Ud&Vd&Wd&Xd&Yd&Zd&ae&be&ce&de&ee&fe&ge&he&ie&j
e&ke&le&me&ne&oe&pe&qe&re&se&te&ue&ve&we&xe&ye&ze&Ae&Be&Ce&De&Ee&Fe&Ge&He&I
e&Je&Ke&Le&Me&Ne&Oe&Pe&Qe&Re&Se&Te&Ue&Ve&We&Xe&Ye&Ze&af&bf&cf&df&ef&ff&gf&h
f&if&jf&kf&lf&mf&nf&of&pf&qf&rf&sf&tf&uf&vf&wf&xf&yf&zf&Af&Bf&Cf&Df&Ef&Ff&G
f&Hf&If&Jf&Kf&Lf&Mf&Nf&Of&Pf&Qf&Rf&Sf&Tf&Uf&Vf&Wf&Xf&Yf&Zf&ag&bg&cg&dg&eg&f
g&gg&hg&ig&jg&kg&lg&mg&ng&og&pg&qg&rg&sg&tg&ug&vg&wg&xg&yg&zg&Ag&Bg&Cg&Dg&E
g&Fg&Gg&Hg&Ig&Jg&Kg&Lg&Mg&Ng&Og&Pg&Qg&Rg&Sg&Tg&Ug&Vg&Wg&Xg&Yg&Zg&ah&bh&ch&d
h&eh&fh&gh&hh&ih&jh&kh&lh&mh&nh&oh&ph&qh&rh&sh&th&uh&vh&wh&xh&yh&zh&Ah&Bh&C
h&Dh&Eh&Fh&Gh&Hh&Ih&Jh&Kh&Lh&Mh&Nh&Oh&Ph&Qh&Rh&Sh&Th&Uh&Vh&Wh&Xh&Yh&Zh&ai&b
i&ci&di&ei&fi&gi&hi&ii&ji&ki&li&mi&ni&oi&pi&qi&ri&si&ti&ui&vi&wi&xi&yi&zi&A
i&Bi&Ci&Di&Ei&Fi&Gi&Hi&Ii&Ji&Ki&Li&Mi&Ni&Oi&Pi&Qi&Ri&Si&Ti&Ui&Vi&Wi&Xi&Yi&Z
i&aj&bj&cj&dj&ej&fj&gj&hj&ij&jj&kj&lj&mj&nj&oj&pj&qj&rj&sj&tj&uj&vj&wj&xj&y
j&zj&Aj&Bj&Cj&Dj&Ej&Fj&Gj&Hj&Ij&Jj&Kj&Lj&Mj&Nj&Oj&Pj&Qj&Rj&Sj&Tj&Uj&Vj&Wj&X
100
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
Golang 1.25.6 and 1.24.12 were released last week. They fix a significant memory exhaustion issue in the parsing of URL-encoded request bodies (CVE-2025-61726). In my testing, a 10 MB request body can result in 416 MB of allocation. A significant risk for any server. Let's dive in, a quick 🧵
100
Reposted by Damien Tournoud
Altmetric @altmetric.com · 08/01/2026
Bluesky is definitely a place for science, research and citing papers. We hope it will continue to close the gap as rapidly as it has with legacy social media. Research, science and dank memes need many homes on the internet. Bluesky provides one of the more inviting ones. Happy New Year.
media.tenor.com
a man wearing a beanie says " yeah science "
Alt: Jeffie (ok Jessie) Pinkman from Breaking Bad wearing a beanie says " yeah science" and points. Hey did anyone watch Pluribus? Sick show, we Stan Vince Gilligan. I bet his middle name is Jeff.
425039
Reposted by Damien Tournoud
Filippo Valsorda @filippo.abyssdomain.expert · 09/01/2026
Here's a fun @standard.site prototype: Atom feeds for any publication! e.g., @octet-stream.net's: geomys-atsite.exe.xyz/profile/octe... Two very cool things: 1. this was ~550 lines of code github.com/FiloSottile/... 2. the feed URL stays the same even if the handle changes or moves PDS!
geomys-atsite.exe.xyz
AT Protocol sites by octet-stream.net
3599