AI & ML interests

The AI community building the future.

Recent Activity

merveย  updated a dataset about 4 hours ago
huggingface/documentation-images
sayakpaulย  updated a dataset about 7 hours ago
huggingface/diffusers-metadata
nielsrย  updated a Space about 11 hours ago
huggingface/paperswithcode
View all activity

Articles

sergiopaniegoย 
posted an update about 6 hours ago
view post
Post
26
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"

you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced

and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine

the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL

blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
  • 2 replies
ยท
nielsrย 
updated a bucket about 18 hours ago
tarekziadeย 
updated a bucket about 21 hours ago
sergiopaniegoย 
posted an update 1 day ago
view post
Post
1313
LFM2.5-2.6B just dropped!

and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.

basically, a full agent training pipeline but compressed into 2.6B

base model โ†’ SFT โ†’ specialized teachers per domain (SFT + RLVR) โ†’ on-policy distillation back into one student โ†’ agentic RL

the two most interesting stages

โ†’ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution

โ†’ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box

this makes a 2.6B that beats much larger models on instruction following and tool use

SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)

โ†’ model: LiquidAI/LFM2.5-2.6B
โ†’ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
โ†’ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
  • 3 replies
ยท
sergiopaniegoย 
posted an update 6 days ago
view post
Post
2573
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now

you look at the drawing and you know. but there is no number, so nothing can train against it, no?

I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL

read the details!๐Ÿค“

https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
  • 2 replies
ยท
sergiopaniegoย 
posted an update 7 days ago
sergiopaniegoย 
posted an update 9 days ago
view post
Post
2871
quick reminder! ๐Ÿšจ

tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series

๐Ÿง  what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
๐Ÿ—“๏ธ when: Tuesday, July 28 - ๐Ÿ•” 5:00 PM CEST / 8:30 PM IST
๐Ÿ“ where: Live on @huggingface 's X, YouTube, and LinkedIn

live: https://www.youtube.com/watch?v=ztdTed5egrM

class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
  • 1 reply
ยท
sergiopaniegoย 
posted an update 12 days ago
view post
Post
200
you can now train your own coding agents with trl + openenv, starting with opencode

we just added end-to-end support for training agent harnesses:

> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs

you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced

we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.

> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv

and we're working actively on both sides so expect more ๐Ÿค“
  • 1 reply
ยท
badaouiย 
posted an update 12 days ago
view post
Post
2314
432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU ๐Ÿคฏ.

Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure ๐Ÿค— Transformers works on day one.

Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.

The result:
โœ… 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.

The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3ร— the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.

A huge thanks to the AMD team for the early access and the great collaboration!

Read the full blog ๐Ÿ‘‡
https://huggingface.co/blog/badaoui/transformers-on-amd-mi455
  • 1 reply
ยท
sergiopaniegoย 
posted an update 13 days ago
sergiopaniegoย 
posted an update 14 days ago
view post
Post
231
join us next Tuesday, July 28, for Class 3 of the Training Agents live series!

we'll dive into reinforcement learning for agent training, covering the intuition behind GRPO, how it works, and how to implement it in TRL with practical, e2e examples

see you there ๐Ÿค 

live: https://www.youtube.com/live/ztdTed5egrM

> in case you missed class 1:
https://x.com/SergioPaniego/status/2069382207618379813
> and in case you missed class 2: https://x.com/SergioPaniego/status/2075180665184686187
evalstateย 
posted an update 26 days ago
view post
Post
1332
Hugging Face MCP Server v0.3.29
~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Included "papers" in the new hf_fs tool. Includes listing of trending/daily.

This is a new tool under observation - disable the "Paper Semantic Search" tool for best results.

hf://papers/
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ daily/
โ”‚   โ”œโ”€โ”€ latest
โ”‚   โ””โ”€โ”€ YYYY/
โ”‚       โ””โ”€โ”€ MM/
โ”‚           โ””โ”€โ”€ DD/
โ”œโ”€โ”€ trending/
โ””โ”€โ”€ ARXIV_ID/
    โ”œโ”€โ”€ metadata.json
    โ”œโ”€โ”€ paper.md
    โ”œโ”€โ”€ models/
    โ”œโ”€โ”€ datasets/
    โ””โ”€โ”€ spaces/

sergiopaniegoย 
posted an update 28 days ago
view post
Post
7737
Frontier models use distillation as a step of their post-training pipelines.

In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.

I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026

It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL!
  • 3 replies
ยท
albertvillanovaย 
posted an update about 1 month ago
view post
Post
3658
๐ŸŽ‰ KTO is now part of the stable TRL API

As of Promote KTO to stable API, KTOTrainer and KTOConfig have graduated from trl.experimental to the stable trl API. https://github.com/huggingface/trl/pull/6175

This one closes out a long road. Over the past 6+ months, the "Align KTO with DPO" effort landed ~90 PRs methodically bringing KTO up to the standard we hold for stable trainers, one carefully-scoped change at a time:
- Feature parity with DPO: full VLM support (incl. multi-image), sync_ref_model, PEFT + Liger, ZeRO-3 + PEFT dtype fix, pad_to_multiple_of, activation offloading, IterableDataset and dict eval_dataset, remove_unused_columns, and reference-logprob precomputation at init.
- Consistency with DPO: aligned method order and signatures, tokenization, _prepare_dataset, PEFT handling, ref-model preparation for distributed training, and config layout โ€” plus a new DataCollatorForKTO and output format. Metrics moved into _compute_loss and simplified to direct averages via the shared _metrics attribute.
- Removing legacy baggage: dropped encoder-decoder support, BOS/EOS handling, null_ref_context, generate_during_eval, model_init, preprocess_logits_for_metrics, model/ref adapter names, and several dead config knobs.
- Coverage: a full test suite mirroring DPO, text collator tests, VLM tests, and slow tests.
- The promotion itself: the experimental โ†’ stable move (#6175) and shim cleanup (#6287), handled so downstream users get a clean deprecation path.

Honestly, this has been one of the more complex tasks I've taken on since joining the team, not because any single change was hard, but because it demanded sustained consistency across a ~2,000-line trainer, with every branch, comment, and edge case kept in lockstep with DPO.

Huge thanks to everyone who reviewed along the way (especially @qgallouedec ), the incremental review cadence is exactly what kept this maintainable.

KTO now sits on equal footing with our other flagship trainers. ๐Ÿš€
  • 2 replies
ยท
danieldkย 
posted an update about 1 month ago
view post
Post
192
We have recently added Torch Stable ABI support to kernels and kernel-builder. This allows kernel developers to target a particular Torch version and the kernel will be supported on that Torch version and later Torch versions (up to ~2 years).

This makes it much easier to write kernels with long-term support and not just the last two Torch releases.

We have also started rolling out Stable ABI support to kernels in kernels-community, starting with Flash Attention 3, supporting Torch 2.9 and later as well as CUDA versions starting at 12.6:

https://huggingface.co/kernels/kernels-community/flash-attn3/tree/v1/build
  • 1 reply
ยท