Adam Fisch، Shubhendu Trivedi، Fantine Huot، William W. Cohen و 4 نفر دیگر
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lo…
Fengqing Jiang، Yite Wang، Boyi Liu، Zhaoyang Wang و 4 نفر دیگر
Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as mat…
آیا یک مدل زبانی خود را بهبود داده است یا خیر؟ به طور تدریجی، این مدل بر اساس دقت متوسط عمل نمیکند؛ بلکه در مورد نحوه مواجهه با مسائل فردی و نقاط ضعف آن مشکل دارد. ردیابی این تغییرات به معنای تفریق دو برآورد نویزآمیز اس…
1. بهبود جریان و خوانایی
2. ساختار جملهای طبیعیتر در فارسی
3. ثبات در اصطلاحات
4. تصحیح خطاهای ترجمه
مقررات:
- حفظ اصطلاحات و فرمولهای فنی
- حفظ ساختار بخشها
- حفظ استنادها و منابع
فقط متن فارسی بهبود یافته را ارائ…
Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that ret…
Bo Liu، Simon Yu، Yiding Jiang، Ao Qu و 14 نفر دیگر
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier)…
Jayjun Lee، Jessica Yin، Asif Rana، Nicholas Blauch و 6 نفر دیگر
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that…
Zhu Zhang، Jixun Wang، Xiaoang Xu، Xiaorong Wang و 5 نفر دیگر
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible respons…
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a fr…