Install
RedditVote46FlipShareTweet46 SharesRedditVote46FlipShareTweet46 Shares
- 1,173articles · 365d
- 10+ hour agolatest article
- Sep 14, 2025earliest in window
- 97%with images · 26 videos
- 339avg words
- Science & Technology 1,151
- Software Dev. 1,000
- Computers & Electronics 897
- Science & Nature 146
- Jobs & Education 100
- STEM 82
- News 34
- Business & Industrial 17
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
How to Build a Lightweight Vision-Language-Action-Inspired Embodied Agent with Latent World Modeling
4+ mon, 2+ week ago (254+ words) We define the compact Vision-Language-Action-inspired world model....
AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation
1+ mon, 1+ day ago (691+ words) Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework....The post AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation...In this tutorial, we build an end-to-end post-training pipeline for a…...
Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training Workflow
4+ mon, 3+ week ago (883+ words) Hugging Face has released ml-intern, an open-source AI agent designed to automate end-to-end post-training...What […] The post Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training...Hugging Face has released ml-intern, an open-source AI…...
A Coding Guide on LLM Post Training with TRL from Supervised Fine Tuning to DPO and GRPO Reasoning
4+ mon, 1+ week ago (575+ words) In this tutorial, we walk through a complete, hands-on journey of post-training large language models...Also, we […] The post A Coding Guide on LLM Post Training with TRL from Supervised Fine Tuning to DPO...In this tutorial, we walk…...
Nous Research Releases NousCoder-14B: A Competitive Olympiad Programming Model Post-Trained on Qwen3-
7+ mon, 3+ week ago (523+ words) When an inference worker finishes a generation, it sends the completion to a Modal verifier and immediately...
VibeThinker-3B: A 3B Dense Reasoning Model Built on Qwen2.5-Coder-3B With the Spectrum-to-Signal Post-Training
2+ mon, 3+ week ago (749+ words) It is post-trained, not pretrained from scratch....The post-training pipeline runs in four stages....
RA3: Mid-Training with Temporal Action Abstractions for Faster Reinforcement Learning (RL) Post-Training
11+ mon, 5+ day ago (462+ words) new research from Apple, formalizes what “mid-training” should do before reinforcement learning RL post-training...shows mid-training should (1) prune to a compact near-optimal action subspace and (2) shorten […] The post...RA3: Mid-Training with Temporal Action Abstractions for Faster Reinforcement Learning (RL) Post…...
Hugging Face Releases TRL v1.0: A Unified Post-Training Stack for SFT, Reward Modeling, DPO, and GRPO
5+ mon, 1+ week ago (223+ words) In the early stages of the LLM boom, post-training was often treated as an experimental ‘dark art.’...Post-training is the phase where a pre-trained base model is refined to follow instructions, adopt a...
Sigmoidal Scaling Curves Make Reinforcement Learning RL Post-Training Predictable for LLMs
10+ mon, 3+ week ago (186+ words) Why that matters: After ~1–2k GPU-hours, you can fit the curve and forecast whether pushing to 10k–100k GPU-hours is worth it—before you burn the budget. The research also shows power-law fits can produce misleading ceilings unless you only fit at very…...
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode
1+ day, 11+ hour ago (358+ words) SWE-2 builds on the infrastructure and recipe behind SWE-1.7, which was post-trained from Kimi K2.7....