Sign in

Damien Tournoud

@damientournoud.bsky.social
64 followers 18 following 89 posts
PostsRepliesMedia
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
3/ Place two: stripThinking in the gemma4 renderer walks past assistant content and removes `<|channel>...<channel|>` blocks before re-rendering. Comment says it mirrors Hugging Face's `strip_thinking` macro.
model/renderers/gemma4.go:
// stripThinking removes <|channel>...<channel|> thinking blocks from content,
// matching the HF chat template's strip_thinking macro.
func stripThinking(text string) string {
	var result strings.Builder
	for {
		start := strings.Index(text, "<|channel>")
		if start == -1 {
			result.WriteString(text)
			break
		}
		result.WriteString(text[:start])
		end := strings.Index(text[start:], "<channel|>")
		if end == -1 {
			break
		}
		text = text[start+end+len("<channel|>"):]
	}
	return strings.TrimSpace(result.String())
}
100
Damien Tournoud @damientournoud.bsky.social · 28/04/2026
2/ Place one: filterThinkTags in `server/routes.go` walks past assistant messages, runs each through a `<think>`/`</think>` parser, and rewrites `Content` to just the post-thinking remainder. Only fires for qwen3 family and deepseek-r1.
server/routes.go:
func filterThinkTags(msgs []api.Message, m *Model) []api.Message {
	if m.Config.ModelFamily == "qwen3" || model.ParseName(m.Name).Model == "deepseek-r1" {
		finalUserIndex := -1
		for i, msg := range msgs {
			if msg.Role == "user" {
				finalUserIndex = i
			}
		}

		for i, msg := range msgs {
			if msg.Role == "assistant" && i < finalUserIndex {
				// TODO(drifkin): this is from before we added proper thinking support.
				// However, even if thinking is not enabled (and therefore we shouldn't
				// change the user output), we should probably perform this filtering
				// for all thinking models (not just qwen3 & deepseek-r1) since it tends
				// to save tokens and improve quality.
				thinkingState := &thinking.Parser{
					OpeningTag: "<think>",
					ClosingTag: "</think>",
				}
				_, content := thinkingState.AddContent(msg.Content)
				msgs[i].Content = content
			}
		}
	}
	return msgs
}
100
Damien Tournoud @damientournoud.bsky.social · 11/03/2026
What could possibly be happening in that LLM to get this garbage as an output?
Google AI overview for the search "strait of hormuz gps interference"

The model output is all garbled, with commas all over the place, like:

"These COMMA electronic warfare attacks create COMMA dangerous navigation COMMA conditions COMMA leading to COMMA ship collisions COMMA drifting COMMA and COMMA skyrocketing insurance COMMA premiums for COMMA carriers"Continued garbled outputContinue garbled output
000
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
Yup. It can be very quick to prematurely declare Mission Accomplished. I have also seen it randomly decide not to do something and declare victory anyway.
Picture of the George W. Bush's 2003 "Mission Accomplished" speech
000
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
4/ Figuring out which years have messages without listing everything: scan backwards from the newest message. Record its year, jump to the start of that year, find the next oldest. Same for months and days. Empty periods cost zero API calls.
gmail.go:
func (g *GmailClient) PopulatedYears(ctx context.Context, labelID string, cache *Cache) ([]int, error) {
	var years []int
	cursor := time.Now().AddDate(1, 0, 0).Unix()
	for {
		s, found, err := g.newestMessageInRange(ctx, labelID, 0, cursor, cache)
		if err != nil {
			return nil, err
		}

		if !found {
			break
		}

		y := time.UnixMilli(s.InternalDate).In(time.Local).Year()
		years = append(years, y)
		cursor = time.Date(y, 1, 1, 0, 0, 0, 0, time.Local).Unix()
	}

	slices.Reverse(years)
	return years, nil
}
100
Damien Tournoud @damientournoud.bsky.social · 27/02/2026
3/ Every label becomes a folder, every email a .eml file, organized by year/month/day. Only populated directories are shown — no scrolling past empty months.
/mnt/gmail/:
All Mail/
  2024/
    01/
      15/
        2024-01-15T143022-Meeting-notes.eml
        2024-01-15T091205-Your-order-has-shipped.eml
Inbox/
  2026/
    02/
      27/
        2026-02-27T083412-Re-Project-update.eml
        2026-02-27T071055-Weekly-digest.eml
Sent/
  2026/
    02/
      26/
        2026-02-26T154830-Re-Lunch-tomorrow.eml
100
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
This results in the allocation of a map with 2.1 million entries, and the corresponding 2.1 million short key strings. In total, 416 MB of allocation.
Parsing a 10 MB body results in 416 MB of allocation over 2.1 million allocations:
$ go test -v -bench=. -benchmem .
BenchmarkMemory-20             1        1126808685 ns/op
        437066424 B/op   2142306 allocs/op
100
Damien Tournoud @damientournoud.bsky.social · 20/01/2026
The issue is a classic memory amplification issue: while Request.ParseForm refuses by default to parse request bodies bigger than 10 MB, even that can result in significant memory usage for specially crafted inputs. In this case, the input is a URL-encoded form data with small keys and no values.
Example application/x-www-form-urlencoded body:
a&b&c&d&e&f&g&h&i&j&k&l&m&n&o&p&q&r&s&t&u&v&w&x&y&z&A&B&C&D&E&F&G&H&I&J&K&L
&M&N&O&P&Q&R&S&T&U&V&W&X&Y&Z&ab&bb&cb&db&eb&fb&gb&hb&ib&jb&kb&lb&mb&nb&ob&p
b&qb&rb&sb&tb&ub&vb&wb&xb&yb&zb&Ab&Bb&Cb&Db&Eb&Fb&Gb&Hb&Ib&Jb&Kb&Lb&Mb&Nb&O
b&Pb&Qb&Rb&Sb&Tb&Ub&Vb&Wb&Xb&Yb&Zb&ac&bc&cc&dc&ec&fc&gc&hc&ic&jc&kc&lc&mc&n
c&oc&pc&qc&rc&sc&tc&uc&vc&wc&xc&yc&zc&Ac&Bc&Cc&Dc&Ec&Fc&Gc&Hc&Ic&Jc&Kc&Lc&M
c&Nc&Oc&Pc&Qc&Rc&Sc&Tc&Uc&Vc&Wc&Xc&Yc&Zc&ad&bd&cd&dd&ed&fd&gd&hd&id&jd&kd&l
d&md&nd&od&pd&qd&rd&sd&td&ud&vd&wd&xd&yd&zd&Ad&Bd&Cd&Dd&Ed&Fd&Gd&Hd&Id&Jd&K
d&Ld&Md&Nd&Od&Pd&Qd&Rd&Sd&Td&Ud&Vd&Wd&Xd&Yd&Zd&ae&be&ce&de&ee&fe&ge&he&ie&j
e&ke&le&me&ne&oe&pe&qe&re&se&te&ue&ve&we&xe&ye&ze&Ae&Be&Ce&De&Ee&Fe&Ge&He&I
e&Je&Ke&Le&Me&Ne&Oe&Pe&Qe&Re&Se&Te&Ue&Ve&We&Xe&Ye&Ze&af&bf&cf&df&ef&ff&gf&h
f&if&jf&kf&lf&mf&nf&of&pf&qf&rf&sf&tf&uf&vf&wf&xf&yf&zf&Af&Bf&Cf&Df&Ef&Ff&G
f&Hf&If&Jf&Kf&Lf&Mf&Nf&Of&Pf&Qf&Rf&Sf&Tf&Uf&Vf&Wf&Xf&Yf&Zf&ag&bg&cg&dg&eg&f
g&gg&hg&ig&jg&kg&lg&mg&ng&og&pg&qg&rg&sg&tg&ug&vg&wg&xg&yg&zg&Ag&Bg&Cg&Dg&E
g&Fg&Gg&Hg&Ig&Jg&Kg&Lg&Mg&Ng&Og&Pg&Qg&Rg&Sg&Tg&Ug&Vg&Wg&Xg&Yg&Zg&ah&bh&ch&d
h&eh&fh&gh&hh&ih&jh&kh&lh&mh&nh&oh&ph&qh&rh&sh&th&uh&vh&wh&xh&yh&zh&Ah&Bh&C
h&Dh&Eh&Fh&Gh&Hh&Ih&Jh&Kh&Lh&Mh&Nh&Oh&Ph&Qh&Rh&Sh&Th&Uh&Vh&Wh&Xh&Yh&Zh&ai&b
i&ci&di&ei&fi&gi&hi&ii&ji&ki&li&mi&ni&oi&pi&qi&ri&si&ti&ui&vi&wi&xi&yi&zi&A
i&Bi&Ci&Di&Ei&Fi&Gi&Hi&Ii&Ji&Ki&Li&Mi&Ni&Oi&Pi&Qi&Ri&Si&Ti&Ui&Vi&Wi&Xi&Yi&Z
i&aj&bj&cj&dj&ej&fj&gj&hj&ij&jj&kj&lj&mj&nj&oj&pj&qj&rj&sj&tj&uj&vj&wj&xj&y
j&zj&Aj&Bj&Cj&Dj&Ej&Fj&Gj&Hj&Ij&Jj&Kj&Lj&Mj&Nj&Oj&Pj&Qj&Rj&Sj&Tj&Uj&Vj&Wj&X
100
Damien Tournoud @damientournoud.bsky.social · 06/01/2026
5/ That's not the only choice you can make. As noted in the Sync 1.1 release notes, if the repository was ordered in preorder depth-first search order, a reader implementation could both validate the cryptographic properties and iterate the keys in order.
A screenshot of https://docs.bsky.app/blog/relay-sync-updates that reads:
    
    The need for ordered repository CAR file exports has become more clear, and an early implementation was completed for the PDS reference implementation. That implementation is not performant enough to merge yet, and it may be some time before ordered CAR files are a norm in the network. The exact ordering also needs to be described more formally to ensure interoperation. Work has not yet started on the "partial synchronization" variant of getRepo, which will allow fetching a subset of the repository.
100
Damien Tournoud @damientournoud.bsky.social · 05/01/2026
5/ Tap itself fetches the list of known repositories from the relay via the com.atproto.sync.listRepos method. A full fetch of the list takes a few hours, because the API scales pretty poorly. This is what the rate of discovery of new repositories looked like the first time I ran this:
A graph of the rate of discovery of new repositories (in iops/s) over time. It shows the rate plummeting from 1.3 kops/s at start to 250 ops/s a few hours later.
110
Damien Tournoud @damientournoud.bsky.social · 31/12/2025
10/ The biggest collection excluding Bluesky (jp.5leaf.sync.mastodon) is only 280MB compressed in total. On the tail end, there is a large amount of very small collections, most of them appearing spamish.
A bar chart of the size of the biggest collections on the ATProto network, excluding the Bluesky collections.
242
Damien Tournoud @damientournoud.bsky.social · 31/12/2025
9/ At the collection level, the Bluesky-related collections (app.bsky.* and chat.bsky.*) form the overwhelming bulk of the data stored on the network (literally 99.9% in size). In decreasing order of size: likes (a majority of the data), then posts, then reposts and finally follows.
A bar chart of the size of the biggest collections on the ATProto network.
142
Damien Tournoud @damientournoud.bsky.social · 31/12/2025
8/ Speaking of compression, in my current storage (CARv1 format, zstd level 1), most repositories achieve a compression ratio above 2x. Still, 20% of the repositories compress very little, at 1.51x ratio. Surprising given the amount of redundancy in the data (i.e. structs are stored as maps).
A graph of the distribution of the compression ratio of all the ATProto repositories synchronized by tap.
241
Damien Tournoud @damientournoud.bsky.social · 31/12/2025
7/ On the other side of the size distribution, the biggest repository is 457MB compressed, and 1.18GB uncompressed. Definitely some interesting outliers here.
A table of some of the biggest repositories on the ATProto network.
162
Damien Tournoud @damientournoud.bsky.social · 31/12/2025
6/ Not unexpectedly, most of the repositories are very small. 20% of the repositories are below 912B while 50% of the repositories are below 6.34kB. You have to wait until the 97th centile to reach 1.04MB.
A graph of the distribution of the compressed size of all the ATProto repositories synchronized by tap.
1113
Damien Tournoud @damientournoud.bsky.social · 04/11/2025
GitHub, my boy. Dates are hard.
A screenshot of the billing page of GitHub.com.

It shows a warning banner with the text "Your most recent billing attempt failed. We will try again on July 11, 2026."

This is might be a date parsing issue, somewhere something parsed November 7th (7/11/2025) as July 11th. But how did it get 2026 for the year?
010
Damien Tournoud @damientournoud.bsky.social · 16/09/2025
What you are looking at here is the 301 response of Nginx 1.22.1.
A censys.io screenshot showing that the body hash c389e590871a87f27ad27393cf7f2947c3ede6ba1cca818cbcff4131e0d0eac4 is a 301 response page generated by Nginx 1.22.1.
000