🔄 In a Training Loop
Ash W.
Hoglet-33
AI & ML interests
Open source AI, datasets, parameter efficiency, SLMs, AI for the betterment of humanity. Contact at ash@basicallyai.co
Recent Activity
reacted to Banaxi-Tech's post with 😎 about 18 hours ago
We're releasing Overfitter 1.0.
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
https://huggingface.co/BananaMind/Overfitter-1.0 upvoted an article 1 day ago
Extreme Overtraining in Tiny Language Models repliedto Banaxi-Tech's post 4 days ago
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
https://huggingface.co/BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
https://huggingface.co/bananamind-research-community