Selected work
stampede
Collapsedodo_v1, 14.5M samples
81%of episodes end in a fall. The first policy, 14.5M samples in: nineteen in a hundred stay up, and those drift backwards.
Shufflegpu2, 2.67B samples
0.07 m/sforward, 2.67 billion samples later. 85% stay up now. They shuffle; they do not stride.
Stridegpu6, 567M samples
59%of the commanded speed, tracked. The best walker, 567M samples with a gait clock and an imitation term in the recipe: 85% stay up, and it is a gait.
4,096robots learning at once.
430kcontrol steps per second, on one rented RTX 3090.
No physical robot walks yet. Policies fall on balance within a minute, and getting up after a fall is still at zero.
0 of 9adversarial breaches with the barrier filter on, 9 of 9 with it off
muzzle and imp
A hardware firewall on a robot's servo bus: the brain proposes, a few-dollar chip disposes.
1,746 to 60cache-write tokens per reminder turn, before and after the LiteLLM fix
Cache keys that forget a dimension
Fixes in LiteLLM, TensorRT-LLM and vLLM for prompt and KV caches that silently stop hitting.
5 to 9xmore tokens per Bengali word than a dedicated tokenizer
What Bengali costs a tokenizer
A paper and an open harness measuring how much more Bengali text costs under multilingual tokenizers.
1 in 1,000episodes where a different build of the engine changed whether the robot fell
Does the compiler decide if the robot falls?
A pre-registered test of whether the build of the physics engine changes a policy's outcomes.
Now
Training a small biped across thousands of simulated copies at once, and building muzzle, the servo-bus layer for a robot that cannot be talked out of its safety limits.
Fixed prompt-cache and KV-cache keys upstream: two fixes merged in LiteLLM, open PRs in TensorRT-LLM and vLLM. Started imp.
Engineering intern at The Breadwinners Club.
Measured what Bengali costs a tokenizer, and traced most of the waste to one character.
Notes
What I believed, and what the measurement said instead.
The robot was reading a sensor that does not exist
The policy's orientation input was the body axis in the world frame. No IMU can measure that, and it sat 25 to 41 degrees from what a real one reports. Every checkpoint trained before the fix was invalid for transfer.
Every "walks for 20 minutes" was false
The endurance evaluator never counted falls. Re-run correctly, the median walk before a fall is 0.3 to 0.7 minutes, for every policy.
The compiler does not decide if the robot falls
Three builds of the physics engine diverge bitwise in nearly every episode, yet outcomes flip no more often than a one-ulp nudge. Bias above about 1 percent is excluded.
My bug finder found the bug, then nothing new
A static analyzer for cache keys that drop a dimension re-found the TensorRT-LLM defect from a cold start, ranked at the top of what it flagged across 21,442 functions. It missed the vLLM instance, and after nine repositories it has found no new defect.
About
I'm Shifat Santo. I study computer science at the University of Texas at Dallas and finish in August 2027.
I work where learning meets hardware: policies trained across thousands of simulated robots, the bus between a robot's brain and its joints, and the inference stacks that serve the models. I publish the numbers, including the ones that go against me.
