RDTvlokip's picture
๐Ÿ‘‹ Open to Work

RDTvlokip PRO

RDTvlokip

AI & ML interests

I'm builds complete AI systems from the ground up, tokenizers, search engines, and small-scale language models to investigate, through exact mathematics and public experimentation, how language and intelligence emerge from minimal resources and pure signal.

Recent Activity

repliedto their post about 9 hours ago
I published an article about training a network to write from reward alone. Code on GitHub with it. Then someone read the code. Dipankar Sarkar commented four times in a day. Each time he had run something first. He rebuilt my statistics in numpy because he had no torch installed. He found a bound I had missed. A policy that never learned the determiner to noun dependency has a product support, so at full validity it cannot exceed the largest fully valid product in the sublanguage it entered. That is 12 on one side and 24 on the other, computable before any training. Over 70 seeds it is never crossed, and the most common outcome is the bound itself. I had published one of those numbers as an interesting coincidence. Then four of my published numbers came apart. Three were a single seed. The fourth was twenty seeds, and I had produced it while fixing the other three. And the test I built to validate his bound tested nothing. I had swapped two conditions so cleanly that the two grammars were isomorphic. Seventy seeds would have returned the mirror image by construction. The real lesson: A relabelling can permute, but it cannot change a ratio. A perfectly symmetric control is often a perfectly empty one. None of my errors were in the reasoning. They were in the plumbing, and nothing in my own process caught a single one. Code, figures, and the notebook with eight dated refutations ๐Ÿ‘‡ ๐Ÿ”— https://huggingface.co/blog/RDTvlokip/i-published-my-rl-experiments ๐Ÿ’ป https://github.com/RDTvlokip/RDTRL ๐Ÿ“ฆ https://doi.org/10.5281/zenodo.21726216
repliedto their post 1 day ago
I published a second article about a reader taking apart four of my numbers. He came back a fifth time, and went to the experiment that had never run. He found the one derived threshold in it was built on a sample maximum. He was right, and there was worse. That maximum cannot converge, because the codes I wanted to declare unreachable are themselves inside the null distribution. Its supremum is exactly the value I was excluding. So I ran the whole thing. Seven questions in one day, in a world of 27 referents small enough that the optimum, the null distribution and the gradient are computed rather than estimated. The certificate carrying the project does not survive two agents. A symmetry argument replaces it, with a corollary I did not expect: a free per-object embedding table cancels in advance anything the message structure could contribute. No training trick recovers it. You can check that by reading an architecture, without running it. Eight of my hypotheses died that day. Not one was an arithmetic error. Then I did the literature review, last, and found the argument published in 2021. The real lesson: Shrinking a world until everything is exact removes one class of mistake and leaves untouched the class that was doing the damage. Twenty minutes of searching would have saved a day of deriving. What caught my errors was never the exactness. It was an outside reader, a second route to the same number, and predictions written down before measuring. Code, every number including the ones I would have cut, and sixteen dated refutations ๐Ÿ‘‡ ๐Ÿ”— https://huggingface.co/blog/RDTvlokip/i-made-my-world-small-enough-to-compute-everything ๐Ÿ’ป https://github.com/RDTvlokip/RDTRL ๐Ÿ“ฆ https://doi.org/10.5281/zenodo.21726216
repliedto their post 1 day ago
I published an article about training a network to write from reward alone. Code on GitHub with it. Then someone read the code. Dipankar Sarkar commented four times in a day. Each time he had run something first. He rebuilt my statistics in numpy because he had no torch installed. He found a bound I had missed. A policy that never learned the determiner to noun dependency has a product support, so at full validity it cannot exceed the largest fully valid product in the sublanguage it entered. That is 12 on one side and 24 on the other, computable before any training. Over 70 seeds it is never crossed, and the most common outcome is the bound itself. I had published one of those numbers as an interesting coincidence. Then four of my published numbers came apart. Three were a single seed. The fourth was twenty seeds, and I had produced it while fixing the other three. And the test I built to validate his bound tested nothing. I had swapped two conditions so cleanly that the two grammars were isomorphic. Seventy seeds would have returned the mirror image by construction. The real lesson: A relabelling can permute, but it cannot change a ratio. A perfectly symmetric control is often a perfectly empty one. None of my errors were in the reasoning. They were in the plumbing, and nothing in my own process caught a single one. Code, figures, and the notebook with eight dated refutations ๐Ÿ‘‡ ๐Ÿ”— https://huggingface.co/blog/RDTvlokip/i-published-my-rl-experiments ๐Ÿ’ป https://github.com/RDTvlokip/RDTRL ๐Ÿ“ฆ https://doi.org/10.5281/zenodo.21726216
View all activity

Organizations

Blog-explorers's profile picture Hugging Face Discord Community's profile picture