🤝 Open to Collab
Shyam Sunder Kumar
theainerd
AI & ML interests
Natural Language Processing
Recent Activity
reacted to Banaxi-Tech's post with ❤️ 2 days ago
AGI has arrived.
Just gotta wait for the GLM distill. upvoted a collection 4 days ago
NASA-IBM-Lunar-FM and Downstream Models liked a model 4 days ago
deepseek-ai/DeepSeek-V4.1-FlashOrganizations
reacted to Banaxi-Tech's post with ❤️ 2 days ago
reacted to prithivMLmods's post with 🔥 16 days ago
Post
3081
ImageShield-MMCF — Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.
⊹ ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
⊹ ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.
⊹ ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
⊹ ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
reacted to FredyRivera-dev's post with ❤️ about 2 months ago
Post
6407
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen.
Covers:
- Hybrid architecture based on Qwen3.5
- Pre-training with 15B tokens
- Cost benchmark between H200 and B200
- Post-training with SFT + LoRA
- Full code and data, open source
With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline.
Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch
Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
Covers:
- Hybrid architecture based on Qwen3.5
- Pre-training with 15B tokens
- Cost benchmark between H200 and B200
- Post-training with SFT + LoRA
- Full code and data, open source
With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline.
Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch
Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
reacted to Tonic's post with 🔥 7 months ago
Post
3877
🤔 Who would win ?
- a fully subsidized ai lab
- 3 random students named
kurakurai ?
demo : Tonic/fr-on-device
if you like it give the demo a little star and send a shoutout to : @MaxLSB @jddqd and @GAD-cell for absolutely obliterating the pareto frontier of the french language understanding .
- a fully subsidized ai lab
OR - 3 random students named
demo : Tonic/fr-on-device
if you like it give the demo a little star and send a shoutout to : @MaxLSB @jddqd and @GAD-cell for absolutely obliterating the pareto frontier of the french language understanding .
reacted to danielhanchen's post with 🚀🚀 7 months ago
Post
2745
We collabed with HF on showing how you can use HF Jobs and Unsloth! https://huggingface.co/blog/unsloth-jobs
reacted to AdinaY's post with 🚀🚀 7 months ago
Post
4581
MiniMax M2.5 is now available on the hub 🚀
MiniMaxAI/MiniMax-M2.5
✨ 229B - Modified MIT license
✨37% faster than M2.1
✨ ~$1/hour at 100 TPS
MiniMaxAI/MiniMax-M2.5
✨ 229B - Modified MIT license
✨37% faster than M2.1
✨ ~$1/hour at 100 TPS
reacted to AdinaY's post with ❤️ 8 months ago
Post
2698
From ChatGPT Healthcare to Claude for healthcare, AI in medicine is speeding up🚀
Now BaichuanAI joins with Baichuan-M3 🏥 an open medical LLM trained for clinical decision-making
https://huggingface.co/collections/baichuan-inc/baichuan-m3
✨ 235B - Apache2.0
✨ Lower hallucinations via Fact-Aware RL
✨ Built for long medical chats
Now BaichuanAI joins with Baichuan-M3 🏥 an open medical LLM trained for clinical decision-making
https://huggingface.co/collections/baichuan-inc/baichuan-m3
✨ 235B - Apache2.0
✨ Lower hallucinations via Fact-Aware RL
✨ Built for long medical chats
reacted to AdinaY's post with 🚀 8 months ago
Post
2886
AgentCPM-Explore🔥 on device agent foundation model released by OpenBMB
openbmb/AgentCPM-Explore
✨ 4B - Apache2.0
✨ Supports 100+ multi-turn environment interactions with search + verification
✨ Full training/inference stack is openly shared as well
openbmb/AgentCPM-Explore
✨ 4B - Apache2.0
✨ Supports 100+ multi-turn environment interactions with search + verification
✨ Full training/inference stack is openly shared as well
reacted to danielhanchen's post with 🔥 10 months ago
Post
8608
Qwen3-Next can now be Run locally! (30GB RAM)
Instruct GGUF: unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF
The models come in Thinking and Instruct versions and utilize a new architecture, allowing it to have ~10x faster inference than Qwen32B.
💜 Step-by-step Guide: https://docs.unsloth.ai/models/qwen3-next
Thinking GGUF: unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF
Instruct GGUF: unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF
The models come in Thinking and Instruct versions and utilize a new architecture, allowing it to have ~10x faster inference than Qwen32B.
💜 Step-by-step Guide: https://docs.unsloth.ai/models/qwen3-next
Thinking GGUF: unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF
reacted to mitkox's post with 🔥 10 months ago
Post
3257
I run 20 AI coding agents locally on my desktop workstation at 400+ tokens/sec with MiniMax-M2. It’s a Sonnet drop-in replacement in my Cursor, Claude Code, Droid, Kilo and Cline peak at 11k tok/sec input and 433 tok/s output, can generate 1B+ tok/m.All with 196k context window. I'm running it for 6 days now with this config.
Today max performance was stable at 490.2 tokens/sec across 48 concurrent clients and MiniMax M2.
Z8 Fury G5, Xeon 3455, 4xA6K. Aibrix 0.5.0, vLLM 0.11.2,
Today max performance was stable at 490.2 tokens/sec across 48 concurrent clients and MiniMax M2.
Z8 Fury G5, Xeon 3455, 4xA6K. Aibrix 0.5.0, vLLM 0.11.2,
reacted to sergiopaniego's post with 👍 10 months ago
Post
2618
we've just added several example scripts to TRL showing how to train models with GRPO using some of the new OpenEnv environments
train a model to interact with a browser (🎮 BrowserGym Env), play Wordle (🎮 Wordle Env) and moooore!
TRL (GRPO + vLLM) + OpenEnv! ⚡️
📝 go play with them: https://github.com/huggingface/trl/tree/main/examples/scripts/openenv
📝 examples list: https://huggingface.co/docs/trl/main/en/example_overview#scripts
train a model to interact with a browser (🎮 BrowserGym Env), play Wordle (🎮 Wordle Env) and moooore!
TRL (GRPO + vLLM) + OpenEnv! ⚡️
📝 go play with them: https://github.com/huggingface/trl/tree/main/examples/scripts/openenv
📝 examples list: https://huggingface.co/docs/trl/main/en/example_overview#scripts
reacted to flozi00's post with 👍 10 months ago
Post
2788
Running large language models efficiently is more than just raw GPU power. The latest guide breaks down the essential math to determine if your LLM workload is compute-bound or memory-bound.
We apply these principles to a real-world example: Qwen's 32B parameter model on the new NVIDIA RTX PRO 6000 Blackwell Edition.
In this guide, you will learn how to:
Calculate your GPU's operational intensity (Ops:Byte Ratio)
Determine your model's arithmetic intensity
Identify whether your workload is memory-bound or compute-bound
Read the full guide here: https://flozi.net/en/guides/ai/llm-inference-math
We apply these principles to a real-world example: Qwen's 32B parameter model on the new NVIDIA RTX PRO 6000 Blackwell Edition.
In this guide, you will learn how to:
Calculate your GPU's operational intensity (Ops:Byte Ratio)
Determine your model's arithmetic intensity
Identify whether your workload is memory-bound or compute-bound
Read the full guide here: https://flozi.net/en/guides/ai/llm-inference-math
posted an update 10 months ago
Post
5049
Hindi Speech to Text just crossed 20 million downloads. Grateful for everyone using it.
theainerd/Wav2Vec2-large-xlsr-hindi
theainerd/Wav2Vec2-large-xlsr-hindi
reacted to AdinaY's post with 🚀 10 months ago
Post
3405
Kimi K2 Thinking is now live on the hub 🔥
moonshotai/Kimi-K2-Thinking
✨ 1T MoE for deep reasoning & tool use
✨ Native INT4 quantization = 2× faster inference
✨ 256K context window
✨ Modified MIT license
moonshotai/Kimi-K2-Thinking
✨ 1T MoE for deep reasoning & tool use
✨ Native INT4 quantization = 2× faster inference
✨ 256K context window
✨ Modified MIT license
reacted to sondhiArm's post with 🔥 11 months ago
Post
1442
Arm will be @ PyTorch Conference, Join Us!
Join us on site October 22-23 to see how Arm empowers developers to build and deploy AI applications with ease using PyTorch and ExecuTorch. Learn about the latest AI technologies from Arm and our ecosystem while expanding your professional network alongside like-minded AI engineers.
Learn more here:
https://huggingface.co/blog/Arm/arm-at-pytorch-conference
Join us on site October 22-23 to see how Arm empowers developers to build and deploy AI applications with ease using PyTorch and ExecuTorch. Learn about the latest AI technologies from Arm and our ecosystem while expanding your professional network alongside like-minded AI engineers.
Learn more here:
https://huggingface.co/blog/Arm/arm-at-pytorch-conference
reacted to AdinaY's post with 👀 12 months ago
Post
1672
Ring-1T-preview 🔥 1T thinking model released by Ant Group.
https://huggingface.co/inclusionAI/Ring-1T-preview
✨ MoE architecture + 20T tokens + RLVR via ASystem
✨ Strong natural language reasoning (AIME’25: 92.6, close to GPT-5)
✨IMO tests: advanced problem-solving & reasoning
https://huggingface.co/inclusionAI/Ring-1T-preview
✨ MoE architecture + 20T tokens + RLVR via ASystem
✨ Strong natural language reasoning (AIME’25: 92.6, close to GPT-5)
✨IMO tests: advanced problem-solving & reasoning
reacted to hesamation's post with ❤️ about 1 year ago
Post
14921
a senior engineer at google just dropped a 400-page free book on docs for review: agentic design patterns.
the table of contents looks like everything you need to know about agents + code:
> advanced prompt techniques
> multi-agent patterns
> tool use and MCP
> you name it
read it here: https://docs.google.com/document/d/1rsaK53T3Lg5KoGwvf8ukOUvbELRtH-V0LnOIFDxBryE/edit?tab=t.0#heading=h.pxcur8v2qagu
you can also pre-order on Amazon (published by Springer) and the royalties goes to Save the Children: https://www.amazon.com/Agentic-Design-Patterns-Hands-Intelligent/dp/3032014018/
the table of contents looks like everything you need to know about agents + code:
> advanced prompt techniques
> multi-agent patterns
> tool use and MCP
> you name it
read it here: https://docs.google.com/document/d/1rsaK53T3Lg5KoGwvf8ukOUvbELRtH-V0LnOIFDxBryE/edit?tab=t.0#heading=h.pxcur8v2qagu
you can also pre-order on Amazon (published by Springer) and the royalties goes to Save the Children: https://www.amazon.com/Agentic-Design-Patterns-Hands-Intelligent/dp/3032014018/
reacted to danielhanchen's post with ❤️ about 1 year ago
Post
5739
Run OpenAI's new gpt-oss models locally with Unsloth GGUFs! 🔥🦥
20b GGUF: unsloth/gpt-oss-20b-GGUF
120b GGUF: unsloth/gpt-oss-120b-GGUF
Model will run on 14GB RAM for 20b and 66GB for 120b.
20b GGUF: unsloth/gpt-oss-20b-GGUF
120b GGUF: unsloth/gpt-oss-120b-GGUF
Model will run on 14GB RAM for 20b and 66GB for 120b.