Sign in

Weights & Biases

@weightsbiases.bsky.social
833 followers 228 following 13 posts

The AI developer platform.

PostsRepliesMedia
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
Full write-up, all 677 features and how it cleared the gold line utm.io/urkYG
wandb.ai
Tabular ML didn't die. It was just waiting for agents
Learn how our tabular data agent, took #1 on the MLE-bench tabular slice with two golds, three above-median, four-for-four valid submissions. .
000
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
None of this needed a smarter model. A 2018 XGBoost and a 2025 one score about the same. What never scaled was the engineer running 20 feature loops a week. The agent runs 60 overnight for pocket change, and trying things was most of the job anyway.
100
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
it didn't get there by brute force. Nobody told it crystals are periodic. It worked out that raw atom-to-atom distances misread the chemistry across a cell boundary, then rebuilt them as metal-to-oxygen nearest neighbors. 677 features by the end. 😳
100
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
The headline run was a materials-science task. Predict formation energy and bandgap of a transparent conductor from its crystal structure. It landed 1st of 685. OOF score 0.04853 against a 0.05589 gold line. The humans who won it originally spent weeks.
100
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
Sixty cycles later, across four MLE-bench datasets at once, it had two golds and three of four finishes above the Kaggle median. No human in the loop the whole run. One NVIDIA L40 on @CoreWeave and a few dollars.
100
Weights & Biases @weightsbiases.bsky.social · 08/07/2026
Tabular ML didn't die. It was just waiting for agents. We rented one GPU for the price of a coffee and asked an agent to top a Kaggle leaderboard on its own. It came back with two gold medals. Here's the run 👇
100
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
Ready to test these new features? Add your keys, start generating trials, and unlock the full potential of the playground.
080
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
While trials work seamlessly with most models, some like o1 currently don’t support choices. We’re actively working to enable this in the future. Stay tuned!
170
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
The magic doesn’t stop there. Once you pick the best output, you can continue your exploration as if that output is how the model would’ve answered. See how the conversation unfolds from your chosen result!
160
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
Trials really shines when temperature settings are turned up. 🔥 Explore the model’s creativity by generating multiple outputs at once and comparing the diverse responses—all at a glance.
100
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
Choosing the best model output often means iteration. Playground trials will save you time by letting you compare multiple results side-by-side before committing to one. 🕵
100
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
We’ve added trials to the playground! Now, you can generate multiple outputs for the same input. Simply bump up the “Number of Trials” in the settings sidebar to see various options before continuing your work.
100
Weights & Biases @weightsbiases.bsky.social · 18/12/2024
🛠️ New tools in the W&B playground are here! Amazon Nova models & xAI's Grok beta LLM are now available. Add your API keys in team settings to start exploring these powerful models today. 🚀 Here’s what’s new: 👇
240