amos @fasterthanli.me · 16/09/2026looks like we have JEV at home huggingface.co/harshatheg/Q...huggingface.coharshatheg/Qwen-2.5-1B-RLCD · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 6644
fry69 @fry69.dev · 16/09/2026Homepage for "Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty", which seem to have started all this. -> rl-calibration.github.io 191
fry69 @fry69.dev · 16/09/2026arXiv paper from June 2025 -> arxiv.org/abs/2507.16806arxiv.orgBeyond Binary Rewards: Training LMs to Reason About Their UncertaintyWhen language models (LMs) are trained via reinforcement learning (RL) to generate natural language "reasoning chains", their performance improves on a variety of difficult question answering tasks. T... 030