๐ค Open to Collab
LH-Tech AI
LH-Tech-AI
AI & ML interests
Small AI and ML models. Trained by myself. Completely OpenSource. For you. | Reddit: https://www.reddit.com/user/LH-Tech_AI/
Recent Activity
liked a model about 3 hours ago
gdiamos/amx-reasoning-v1-instruct liked a model about 8 hours ago
FWKV/Myosotis-1-base new activity 1 day ago
slmconsortium/README:New member ratification: GGUFGuyOrganizations
reacted to appvoid's post with ๐ 1 day ago
Post
3573
Supra2-100M is out!
Go check it out:
- https://www.reddit.com/r/LocalLLaMA/comments/1velyl9/new_models_supra2100m_base_and_instruct_go_check/
- https://huggingface.co/SupraLabs/Supra2-100M
- SupraLabs/Supra2-100M-Instruct
Give us a like and a follow!!
HAVE FUN ๐ค๐ฅ๐
more coming soon...
Go check it out:
- https://www.reddit.com/r/LocalLLaMA/comments/1velyl9/new_models_supra2100m_base_and_instruct_go_check/
- https://huggingface.co/SupraLabs/Supra2-100M
- SupraLabs/Supra2-100M-Instruct
Give us a like and a follow!!
HAVE FUN ๐ค๐ฅ๐
more coming soon...
reacted to Enderchef's post with ๐ค๐ฅโค๏ธ 3 days ago
Post
3444
๐ Supra2 100M is out, and multiple other SLM orgs are gaining power!
Following takes a press. Please follow:
fromziro
SupraLabs
AxiomicLabs
Following takes a press. Please follow:
reacted to Banaxi-Tech's post with ๐ค๐ค๐๐ 8 days ago
Post
3505
We're delaying BananaMind 2.1!
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
BananaMind
@Banaxi-Tech
---
@vovaRL
@DedeProGames
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
@Banaxi-Tech
---
@vovaRL
@DedeProGames
Please check your discord DMs @Enderchef ๐ญ
reacted to Banaxi-Tech's post with ๐ฅ 13 days ago
Post
2706
We're releasing Overfitter 1.0.
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
reacted to Bc-AI's post with ๐ฅ 16 days ago
Post
3784
New update! We are currently training a few new models now! Our 3rd generation main LLM standard edition is in training right now. We are also training a new LLM line called Tiny Coder around 350~ish M params. Thanks to @Banaxi-Tech for inspiring the architecture with his Bananamind-2.1-unified test model. Thanks to our beta testers: @juiceb0xc0de @ProCreations @Sbui503 @Fishtiks @MUK-IS-GOAT
reacted to Banaxi-Tech's post with ๐ฅ 16 days ago
Post
3740
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term.
Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk โ it needs both outer towers to become semantic.
The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA.
Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included.
This is a research model. It is interesting because you can take it apart.
BananaMind/BananaMind-2.1-Unified
Follow us for more:
BananaMind
@Banaxi-Tech
Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk โ it needs both outer towers to become semantic.
The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA.
Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included.
This is a research model. It is interesting because you can take it apart.
BananaMind/BananaMind-2.1-Unified
Follow us for more:
@Banaxi-Tech
replied to Banaxi-Tech's post 16 days ago
Looks cool! Can't wait to test these!
But we will beat this with our Supra3 family soon (coming in around 4 to 8 weeks)! ๐ฅ๐
Stay tuned! :D
reacted to Banaxi-Tech's post with ๐ฅ 16 days ago
Post
3790
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
bananamind-research-community
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
@Banaxi-Tech
@vovaRL
@DedeProGames
reacted to Enderchef's post with ๐๐๐ค๐ฅ 19 days ago
Post
4096
Do you support the SLM(1M-150M parameter) community?
If so, join the SLM discord(https://discord.gg/BBYaERvvn), and give these orgs some follows!
BananaMind
SupraLabs
fromziro
AxiomicLabs
If so, join the SLM discord(https://discord.gg/BBYaERvvn), and give these orgs some follows!