Thomas Capelle @capetorch.bsky.social · 21/02/2025This was a team effort from @morgymcg.bsky.social , Soumik, @parambharat.bsky.social , Agata Mlynarczyk, @ayshthkr.bsky.social and many others! 000
Thomas Capelle @capetorch.bsky.social · 21/02/2025I'm excited to see how the community uses these tools, and I'm looking forward to more innovations in safe and reproducible AI! Check the scorers and Weave here: 👉 wandb.me/weave_scorers 📚 A colab: wandb.me/scorers_colabwandb.meLocal Weave Scorers | W&B WeaveWeave's local scorers are a suite of small language models that run locally on your machine with minimal latency. These models evaluate the safety and quality of your AI system’s inputs, context, and ... 100
Thomas Capelle @capetorch.bsky.social · 21/02/2025A personal highlight was working on the Fluency Scorer powered by AnswerDotAI ModernBERT-base; we hope to move all DeBerta-powered scorers to ModernBert in the next release so we can benefit from the longer context length and training speed! 100
Thomas Capelle @capetorch.bsky.social · 21/02/2025As part of this initiative, we also created comprehensive evaluation datasets, drawing on invaluable contributions from the open-source community. Being a reproducibility-first company, we’ve made the full recipe public, including the scorers, model weights, and the training and evaluation datasets 100
Thomas Capelle @capetorch.bsky.social · 21/02/2025We designed these non-LLM powered scorers to leverage state-of-the-art open source models – from the PleIAI/Celadon toxicity detector to the Vectara hallucination scorer – ensuring that our AI systems are evaluated across multiple dimensions. 100
Thomas Capelle @capetorch.bsky.social · 21/02/2025Over the past few months, my team at Weights & Biases has been hard at work launching Weave Scorers and guardrails. wandb.me/weave_scorers 👇wandb.meLocal Weave Scorers | W&B WeaveWeave's local scorers are a suite of small language models that run locally on your machine with minimal latency. These models evaluate the safety and quality of your AI system’s inputs, context, and ... 100
Thomas Capelle @capetorch.bsky.social · 21/02/2025media.tenor.comDisappointed Cat GIFALT: Disappointed Cat GIF 110
Reposted by Thomas CapelleSara Hooker @sarahooker.bsky.social · 13/02/2025Many people have asked me about the France Action Summit. I think a summit is typically most valuable as a catalyst, not as a solution in itself. But, will share some observations. 24210
Thomas Capelle @capetorch.bsky.social · 13/02/2025It could have been called Gulf of North America 000
Thomas Capelle @capetorch.bsky.social · 12/02/2025I just built a CI to run an Eval of some custom LLM scorers on top of @modal-labs.bsky.social - Great to test against different GPUs - No custom runner neded on github - Fast and nice console outputs =) 020
Thomas Capelle @capetorch.bsky.social · 09/02/2025Samedi tu fais moules et frites, dimanche tu finis les moules dans un rissoto aux moules. 100
Reposted by Thomas CapelleAugust J. Pollak @augustjpollak.bsky.social · 08/02/2025Because in the United States, it’s legal to feed chicken shit to cattle. That’s why. That’s literally the reason www.telegraph.co.uk/global-healt... 259148294217
Thomas Capelle @capetorch.bsky.social · 09/02/2025Pancakes morning with the arrival of the @vendeeglobe.bsky.social 000
Thomas Capelle @capetorch.bsky.social · 06/02/2025We raised this internally! thanks for the info. 120
Thomas Capelle @capetorch.bsky.social · 05/02/2025This budget forcing is really smart. We could do that we prefill on API models no? 100
Thomas Capelle @capetorch.bsky.social · 05/02/2025Share your recently used Slack emojis; they tell much about your work ambiance. I like mine =) 000
Thomas Capelle @capetorch.bsky.social · 29/01/2025I mostly use Claude these days, and it works very well In cursor integration. I grab o1 for more complex stuff and when I have a detailed plan and output. How is the API speed and reliability? 100
Thomas Capelle @capetorch.bsky.social · 29/01/2025I understand the R1 hype but are you switching from Claude/o1 to it? 100
Thomas Capelle @capetorch.bsky.social · 28/01/2025media.tenor.coma close up of a man 's face with his eyes closedALT: a close up of a man 's face with his eyes closed 010
Thomas Capelle @capetorch.bsky.social · 28/01/2025It is not there for me (on the paid cursos sub) 100
Thomas Capelle @capetorch.bsky.social · 27/01/2025Should we buy NSQ: NVDA right now? is it going to tank more? 000
Thomas Capelle @capetorch.bsky.social · 23/01/2025@fintual.bsky.social haciendo IA en Chile con Openai 100
Thomas Capelle @capetorch.bsky.social · 23/01/2025Wow this is a cool resource! Way better than injecting random wikipédia titles. 010
Thomas Capelle @capetorch.bsky.social · 21/01/2025yeah, I think I will need to do some seeding somehow 000
Thomas Capelle @capetorch.bsky.social · 21/01/2025How do you get variance on chatGPT/Claude inference to create dataset samples? The model tends to create basically the same content over and over. I ask for a random topic and it talks almost exclusively about cats 🐱. CC @maziyarpanahi.bsky.social @teknium.bsky.social 210
Thomas Capelle @capetorch.bsky.social · 21/01/2025The W&B Programmer agent is now topping the SWE-bench leaderboard. Shawn Lewis (co-founder of @weightsbiases.bsky.social ) has done amazing work on this. Iterating on a complex system like this wouldn't be possible without W&B Weave. Read more about the solution here: medium.com/@shawnup/the... 020