Los resultados los puedes leer en el artículo, Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems, que he desarrollado bajo la tutorización de Oier López de Lacalle, en arxiv.org/abs/2608.30426
arxiv.org
Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems
Current dialogue systems struggle with dynamic information retrieval, often leading to hallucinations and lower response accuracy. We address this by adapting the ReAct framework for Task-Oriented Dia...