Gemma-4-31B-it safety-repair checkpoints: multilingual safety and cybersecurity selectivity. Gated under LSAL v1.2.
AI & ML interests
Frontier research around Safe and aligned intelligence
Recent Activity
View all activity
Papers
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
Datasets released with CircuitKIT and CuratorKIT.
HH-RLHF-aligned starting points and their safety-drifted SFT checkpoints for Llama-3.1/3.2, Gemma-3, and Qwen3.
-
Lexsi/llama32-3b-hh-rlhf-aligned
Text Generation • 3B • Updated • 12 -
Lexsi/gemma3-4b-code-sft-drift
Image-Text-to-Text • 4B • Updated • 2 -
Lexsi/gemma3-4b-dolly-sft-drift
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/gemma3-4b-gsm8k-sft-drift
Image-Text-to-Text • 4B • Updated • 2
DPO-recovered checkpoints that restore safety to the SafeTune drift checkpoints.
-
Lexsi/gemma3-4b-code-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/gemma3-4b-dolly-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/gemma3-4b-gsm8k-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/llama31-8b-code-dpo-recover
Text Generation • 8B • Updated • 3 • 1
obl-ex 4B models and run logs.
-
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
Paper • 2511.02818 • Published • 15 -
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Paper • 2511.02802 • Published • 16 -
Interpretability as Alignment: Making Internal Understanding a Design Principle
Paper • 2509.08592 • Published -
Interpretability-Aware Pruning for Efficient Medical Image Analysis
Paper • 2507.08330 • Published
Gemma-4-31B-it safety-repair checkpoints: multilingual safety and cybersecurity selectivity. Gated under LSAL v1.2.
DPO-recovered checkpoints that restore safety to the SafeTune drift checkpoints.
-
Lexsi/gemma3-4b-code-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/gemma3-4b-dolly-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/gemma3-4b-gsm8k-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 • 1 -
Lexsi/llama31-8b-code-dpo-recover
Text Generation • 8B • Updated • 3 • 1
Datasets released with CircuitKIT and CuratorKIT.
obl-ex 4B models and run logs.
HH-RLHF-aligned starting points and their safety-drifted SFT checkpoints for Llama-3.1/3.2, Gemma-3, and Qwen3.
-
Lexsi/llama32-3b-hh-rlhf-aligned
Text Generation • 3B • Updated • 12 -
Lexsi/gemma3-4b-code-sft-drift
Image-Text-to-Text • 4B • Updated • 2 -
Lexsi/gemma3-4b-dolly-sft-drift
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/gemma3-4b-gsm8k-sft-drift
Image-Text-to-Text • 4B • Updated • 2
-
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
Paper • 2511.02818 • Published • 15 -
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Paper • 2511.02802 • Published • 16 -
Interpretability as Alignment: Making Internal Understanding a Design Principle
Paper • 2509.08592 • Published -
Interpretability-Aware Pruning for Efficient Medical Image Analysis
Paper • 2507.08330 • Published