我About me
Shubham Jadhav · Researcher at IIT Madras · GPU & Graphics · Chennai, India
I am a graduate researcher in Computer Science and Engineering at IIT Madras. I am happiest where elegant algorithms meet real hardware: I take problems that look impossibly large and make them fast, then faster, until they run in the blink of an eye.
My M.S. thesis takes on the Capacitated Vehicle Routing Problem, the puzzle behind every delivery network. Our solver is built from simple, fast heuristics that run in parallel; on million-scale benchmarks it beats the state of the art on average within seconds, about a tenfold cut in runtime on an ordinary 8-core machine.
Real-time graphics is my favourite kind of magic: an OpenGL scene drawn in C on a bare Win32 window, an ocean simulated with CUDA FFTs, and now this valley, where every hill, leaf and note is written in code. In between, I write GPU kernels, parallel graph algorithms and compilers that turn tensor graphs into fused CUDA.
Away from the keyboard I draw, paint and practise calligraphy, play the tabla and the guitar, and never turn down a game of table tennis or chess.
- Studying
- M.S. by Research, CSE · IIT Madras
- Research
- Vehicle routing at million scale
- Focus
- GPU computing, real-time graphics, compilers
- Off the clock
- Tabla, guitar, calligraphy, chess
技Skills
Strike a training dummy to test each discipline.
Graphics
Real-time rendering
- OpenGL
- WebGL
- GLSL shaders
- three.js
- DirectX
- Win32 SDK
Parallel computing
From GPUs to clusters
Languages
Tools of the trade
- C++
- C
- Python
- TypeScript / JavaScript
- Java
- SQL
Workflow
Shipping things
- Git & GitHub
- Linux
- Docker
- Slurm
- LaTeX
- Blender
Engine craft
Under the hood
- Linear algebra for 3D
- Scene graphs & ECS
- Physics & collision
Systems & AI
Compilers, binaries, models
路Journey
Each turn of the bridge is a step on my path.
-
2015 – 2017
School years
Gadhinglaj and Kolhapur
Class X at Gadhinglaj High School with 94.2%, Class XII at Vivekanand College, Kolhapur, and the Road Safety Patrol platoon, which I led through its drills and traffic-safety drives.
-
2017 – 2021
B.Tech in Computer Science
Savitribai Phule Pune University
Graduated with a CGPA of 9.11 and headed the National Service Scheme unit: blood donation drives, check-dam construction and a seven-day rural development camp in Varoti Bk.
-
2024
M.S. by Research begins
IIT Madras
Joined the Department of Computer Science and Engineering and began my thesis on vehicle routing at million scale, guided by Prof. Rupesh Nasre and Prof. Narayanaswamy N S.
-
2024 – now
Teaching and service
IIT Madras
Teaching assistant for GPU Programming, Blockchain Technology and Natural Language Processing; elected MS representative for CSE; and a coordinator of the ICPC India Online Round (2025, 2026) and the CSE degree ceremony.
-
2026
Compilers and secure AI
Built an ML compiler that lowers tensor graphs into fused CUDA and Triton kernels, and showed how prompt injection hijacks AI agents, and how to stop it.
-
Now
Building this valley
And looking for the next hard problem to make fast.
作Projects
Each banner on the path to the pagoda is something I built.
This explorable 3D portfolio: a procedural valley, a hand-animated panda and a generative soundtrack.
Everything you see and hear is generated in code: terrain, wind-blown grass, bamboo, blossom trees, a lake with caustics, pagodas and the music.
Built with three.js, TypeScript and custom GLSL. It draws close to two million triangles in about two hundred draw calls, and an autopilot and a camera director on Bezier paths film the whole story as a five-minute tour.
Source code ↗
Vehicle Routing at Million Scale (M.S. thesis)
2024 – now
A parallel heuristic solver for the Capacitated Vehicle Routing Problem that beats the state of the art on million-scale instances within seconds.
The problem: the cheapest routes from one depot that visit every customer exactly once without overloading a vehicle. It is NP-hard, and it sits at the core of logistics, where every kilometre saved adds up to large savings in cost and emissions.
Our solver pairs a geometry-aware construction with granular local search and ruin-and-recreate, run in parallel over a size-independent decomposition. On million-scale benchmarks it beats the state of the art within seconds and improves best-known solutions within 600 seconds: about ten times faster on an 8-core machine. Guided by Prof. Rupesh Nasre and Prof. Narayanaswamy N S.
Domain-Specific AI Compiler & CUDA Kernel Engine
2026
An ML compiler that lowers tensor graphs through MLIR and LLVM into fused Triton and CUDA kernels, 1.42× faster than PyTorch end to end.
Custom passes fuse operators, remove dead code and unroll loops; a code generator then emits fused Triton and CUDA C++ kernels with shared-memory tiling and float4 loads, so intermediate results never travel back to DRAM.
Measured with NVIDIA Nsight Compute and perf: a 1.42× end-to-end inference speedup and 38% less VRAM bandwidth than the PyTorch baselines.
A real-time rendering demo: an ocean simulated with CUDA FFTs, procedural terrain and atmospheric scattering, filmed by a Bezier camera.
A demo I co-engineered: a CUDA FFT ocean solver, procedural terrain and atmospheric scattering, finished with god rays, bloom and depth of field, and choreographed by Bezier camera animation.
Watch the demo ↗
Delta-stepping shortest paths on NVIDIA GPUs for ultra-large sparse graphs: 300 million edges without host-device synchronisation.
Static and adaptive delta-stepping over graphs in CSR format. Thrust-driven worklists rescale their buckets on the fly to keep 65,536 work items in flight, and native atomics with careful memory reuse do the rest.
Securing Agentic AI Workflows
2026
How a zero-click prompt injection hijacks an LLM agent, and the defences that stop it: human consent, a dual-LLM inspector and a security gateway.
In a LangChain and Llama 3.1 testbed, poisoned data hijacks an agent's ReAct loop for autonomous exfiltration: the excessive agency risk from the OWASP Top 10 for LLM applications.
The defences: human-in-the-loop consent before high-impact tool calls, a dual-LLM "Inspector" that acts as a semantic firewall, and a proposed security gateway with scoped MCP authentication against delegation-chain spoofing.
An OpenGL demo written in C, with Win32 for its window and hand-built data structures for its scenes.
Watch the demo ↗