Sign in

Horace He

@chhillee.bsky.social
833 followers 60 following 13 posts

@PyTorch "My learning style is Horace twitter threads" - @typedfemale

PostsRepliesMedia
Horace He @chhillee.bsky.social · 06/02/2025
Yep that's right! A very common use-case is for "document masking" (i.e. variable length sequences), and that requires recomputing the mask on every iteration (which isn't "free", but is on the order of microseconds to milliseconds and not seconds).
030
Horace He @chhillee.bsky.social · 02/01/2025
What does "there" mean in this case :)
110
Reposted by Horace He
Yaron Minsky @yminsky.bsky.social · 11/12/2024
@chhillee.bsky.social's talk at Jane Street is now up! youtu.be/139UPjoq7Kw?...
youtu.be
Building Machine Learning Systems for a Trillion Trillion Floating Point Operations
YouTube video by Jane Street
0317
Horace He @chhillee.bsky.social · 03/12/2024
I’ll count it!
020
Reposted by Horace He
Mike Smith @mjjsmith.com · 02/12/2024
Getting different attention masks working for AstroPT (a proto-foundation model for astronomy github.com/Smith42/astr...), so much nicer to do it with Flex Attention vs custom CUDA kernels -- thank you for releasing it to the world 🫡
github.com
GitHub - Smith42/astroPT: Transformer for galaxy images (and general astronomy)
Transformer for galaxy images (and general astronomy) - Smith42/astroPT
041
Horace He @chhillee.bsky.social · 01/12/2024
Kinda interesting to me that the books I obsessively read as an elementary schooler are still some of the most popular series today.
130
Horace He @chhillee.bsky.social · 01/12/2024
I think torch-xla is definitely usable if you don’t want to train anything particularly weird or use unusual parallelism schemes. See this tweet from Saining Xie’s lab on evaluating torchxla vs. Jax for their use case: x.com/tongpetersb/...
x.com
x.com
110
Horace He @chhillee.bsky.social · 01/12/2024
The other nice parts about TPUs is that Google gives much more of them out for free compared to GPUs. Arguably this reflects how much people want to use them, but I think it's been a great boon for the academic labs willing to go through the effort.
200
Horace He @chhillee.bsky.social · 01/12/2024
I judge social networks by how many FlexAttention users I can find on each one, and by that metric, Bluesky is doing pretty good!
1501
Horace He @chhillee.bsky.social · 01/12/2024
! What were you using it for?
110
Horace He @chhillee.bsky.social · 01/12/2024
A lot of PyTorch is about dealing with this stuff nowadays!
030
Horace He @chhillee.bsky.social · 01/12/2024
Out of curiosity, what kind of shapes are you typically looking at?
110
Horace He @chhillee.bsky.social · 01/12/2024
Are they actually using FlexAttention here? I didn't see it in the repo
100
Horace He @chhillee.bsky.social · 25/11/2024
If you'd like to influence what features the PyTorch distributed team work on in torchtitan (e.g. MoE, multimodal, context parallelism, etc.), go made your voices heard here!
github.com
Vote on new features! · pytorch torchtitan · Discussion #693
Hi torchtitanists, Thank you for your interests in torchtitan! Please upvote on what features you would like to see next, and add one if it's not already there. We'll try to prioritize on the most ...
0111
Horace He @chhillee.bsky.social · 23/11/2024
First thought: Seems kinda "FlexAttention-y": bsky.app/profile/sungkim.bsky.socia… Second thought: oh cool, they're already using FlexAttention! it's a nice usage of the `or_masks` and `and_masks` API - I think they do (causal & sliding_window) | (register_mask)
090