I released my own benchmark sets behind emotional-memory, my library for mood-aware recall in LLM agents. The card opens with the limits: the main set is built to favour the method, and the negative results are listed up front, e.g. on LoCoMo AFT scores F1 0.168 against 0.271 for a naive RAG baseline. Use it to check where affect-conditioned retrieval helps and where it doesn't. gianlucamazza/emotional-memory-benchmarks
Most NSFW classifiers break the second an image touches the internet.
They look great on pristine benchmarks, but in the wild, every social platform aggressively recompresses, downsamples, and degrades images. The moment JPEG or WebP compression artifacts show up, confidence collapses and false positives spike.
SafeScan was built to survive actual platform pipelines. Trained on 34,000 images under almost every major social media compression profile using a Vision Transformer backbone (google/vit-base-patch16-224). Instead of blunt binary filtering, it breaks decisions down across 5 clear categories:
β’ safe β’ drawing β’ sexy β’ hentai β’ porn
The result is a moderation model that actually generalizes to real-world internet feeds instead of fragile, uncompressed datasets. Open-weight and available on Hugging Face:
The best result so far is 8 of the 14 playable levels cleared in one continuous run. Two models have cleared eight levels so far: - Qwen/Qwen3.8-Flash-Next - Cloudflare/clef
If you'd like me to try a specific model, please name it in the comments.
You can also try beating the game with a model of your choice. The project with instructions for running the challenge with different models/engines is here: https://github.com/felladrin/ai-plays-cat-goric (PRs welcome!)
And attached is a 2-minute recording of Clef clearing eight levels (the game pauses while the model thinks, so the video is time-warped).
Most deepfake audio detectors are quietly cheating.
They donβt really listen to the speech β they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.
AIRealNet-Audio was built to stop that shortcut. It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.
The result is a detector that actually has to learn the artifacts instead of gaming the feature space.