Sign in

xjdr

@xjdr.bsky.social
1.3K followers 91 following 15 posts

hot takes, linear Algebra, JAX apologist, Raconteur

PostsRepliesMedia
xjdr @xjdr.bsky.social · 26/11/2024
... still me ...
070
xjdr @xjdr.bsky.social · 26/11/2024
I have become radicalized
10540
xjdr @xjdr.bsky.social · 25/11/2024
405B Base (bf16) will withstand the test of time and remain eternal (18 months)
1101
xjdr @xjdr.bsky.social · 25/11/2024
so far the experience has been pretty good here but the default feeds are _terrible_. feels like its going to take a few weeks to whip these feeds into shape with mutes and "show less like these" plus lots of likes. Following feed is good but i need to follow a lot more people
15984
xjdr @xjdr.bsky.social · 25/11/2024
theoretically they should be semantically similar / close in latent space but YMMV based on the model.
020
xjdr @xjdr.bsky.social · 25/11/2024
very interesting work and it reminds me a bit of this paper. Tokenizers and ROPE must die. after samplers, i am on to those next ... arxiv.org/abs/2407.036...
arxiv.org
Improving Self Consistency in LLMs through Probabilistic Tokenization
Prior research has demonstrated noticeable performance gains through the use of probabilistic tokenizations, an approach that involves employing multiple tokenizations of the same input string during ...
97912
xjdr @xjdr.bsky.social · 24/11/2024
good call. this isn't universal advice but it is my general advice for most people for most use cases. I have been very surprised with 1B, specifically with function calling and moderately complex coding tasks. punches well above its weight IMHO
120
xjdr @xjdr.bsky.social · 24/11/2024
i keep forgetting to include this cause i always assume people do this by default. Any time there is an exponent or a norm, you should be working in the highest practical precision
0251
xjdr @xjdr.bsky.social · 24/11/2024
my old recommendation used to be run the largest model you can at Q4, but with L3.2 and Qwen2.5 that has changed. Generally, i now suggest run the largest 5T+ token trained model you can at bf16 (not fp16 unless that is how it was trained). L3.2 1B or Qwen 2.5 1.5B are good
140
xjdr @xjdr.bsky.social · 24/11/2024
the BigVision repo is my current reference impl for gemma and ViT. such an underrated repo @giffmana.bsky.social and team are doing the lord's work github.com/google-resea... github.com/google-resea...
github.com
big_vision/big_vision/models/ppp/gemma.py at main · google-research/big_vision
Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more. - google-research/big_vision
39112
xjdr @xjdr.bsky.social · 24/11/2024
now that people are paying attention again, here is your periodic reminder. Always run in bf16. always apply ROPE and attention softmax at float32 (as shown here) github.com/xjdr-alt/ent...
4777
Reposted by xjdr
Alexander Doria @dorialexander.bsky.social · 24/11/2024
So first version of an ml anon starter pack. go.bsky.app/VgWL5L Kept half-anons (like me and Vic). Not all anime pfp, but generally drawn.
106317
xjdr @xjdr.bsky.social · 24/11/2024
i am willing to be hurt again.
050
xjdr @xjdr.bsky.social · 24/11/2024
very solid list, but i am biased
180
xjdr @xjdr.bsky.social · 24/11/2024
i trying to follow as many of my old moots as possible and new people as i find them. some of y'all changing your pfp is just mean spirited (im lazy and learned people's pfps not names)
8361
xjdr @xjdr.bsky.social · 22/11/2024
Well this looks shockingly professional. I may have to put on a tie to post here
6310