Erik @erikkaum.bsky.social · 06/10/2025We have Nvidia B200s ready to go for you in Hugging Face Inference Endpoints 🔥 I tried them out myself and the performance is amazing. On top of that we just got a fresh batch of H100s as well. At $4.5/hour it's a clear winner in terms of price/perf compared to the A100. 061
Erik @erikkaum.bsky.social · 21/03/2025We just refreshed 🍋 our analytics in @hf.co endpoints. More info below! 183
Erik @erikkaum.bsky.social · 13/03/2025Morning workout at the @hf.co Paris office is imo one of the best perks. 030
Erik @erikkaum.bsky.social · 12/03/2025Gemma 3 is live 🔥 You can deploy it from endpoints directly with an optimally selected hardware and configurations. Give it a try 👇 162
Erik @erikkaum.bsky.social · 21/12/2024today as part of a course, I implemented a program that takes a bit stream like so: 10001001110111101000100111111011 and decodes the intel 8088 assembly from it like: mov si, bx mov bx, di only works on the mov instruction, register to register. code: github.com/ErikKaum/bit...t.cohttps://github.com/ErikKaum/bitbubble 010
Erik @erikkaum.bsky.social · 17/12/2024Ambition is a paradox. You should always aim higher, but that easily becomes a state where you're never satisfied. Just reached 10k MRR. Now there's the next goal of 20k. Sharif has a good talk on this: emotional runway. How do you deal with this paradox? video: www.youtube.com/watch?v=zUnQ...youtube.combefore you give up, give this video a chance.YouTube video by Founders, Inc. 010
Reposted by ErikXuan Son Nguyen @ngxson.hf.co · 27/11/2024Hugging Face inference endpoints now support CPU deployment for llama.cpp 🚀 🚀 Why this is a huge deal? Llama.cpp is well-known for running very well on CPU. If you're running small models like Llama 1B or embedding models, this will definitely save tons of money 💰 💰 3266
Reposted by ErikAndi @andimara.bsky.social · 26/11/2024Let's go! We are releasing SmolVLM, a smol 2B VLM built for on-device inference that outperforms all models at similar GPU RAM usage and tokens throughputs. SmolVLM can be fine-tuned on a Google collab and be run on a laptop! Or process millions of documents with a consumer GPU! 410422
Erik @erikkaum.bsky.social · 26/11/2024Is it just me or does it intuitively align that chat bars are at the bottom of the page and search bars at the top? I've noticed that perplexity positions the question on the top and generates the text below. Is it because they want to position more as a search engine? 010
Erik @erikkaum.bsky.social · 23/11/2024typical engineer writing copy in plain english i'd say "2 conversions at the same time" 040
Erik @erikkaum.bsky.social · 22/11/2024lesson: if you care about the performance of something, you gotta run your own benchmarks 0104
Erik @erikkaum.bsky.social · 19/11/2024Just wrote some golang for fun. Damn, I had almost forgotten how enjoyable it’s to program in. Just breezing through the code. If I need thousands of threads, it’s just there. 0120
Erik @erikkaum.bsky.social · 18/11/2024A while ago I started experimenting with compiling the Python interpreter to WASM. To build a secure, fast, and lightweight sandbox for code execution — ideal for running LLM-generated Python code. - Send code simply as a POST request - 1-2ms startup times github.com/ErikKaum/run...github.comGitHub - ErikKaum/runner: Experimental wasm32-unknown-wasi runtime for Python code executionExperimental wasm32-unknown-wasi runtime for Python code execution - ErikKaum/runner 042
Erik @erikkaum.bsky.social · 16/11/2024There are now /llms.txt files for a few of @huggingface.bsky.social docs 🔥 huggingface-projects-docs-llms-txt.hf.space/transformers...huggingface-projects-docs-llms-txt.hf.space 010
Erik @erikkaum.bsky.social · 01/11/2024"If you're thinking without writing, you only think you're thinking." Same applies imo to coding and why it's so important to open your editor, start tinkering and sketching things out. quote from: paulgraham.com/writes.htmlpaulgraham.comWrites and Write-Nots 020