Experience
Part of AML frameworks team. Focused on ML infra which is the core part of recommendation, ads and search direction. Optimised Sparse MoE by +23.3%, increasing throughput of 8B fine ranking model by +15.36%. Optimised flash_attn_varlen_backward, increasing throughput by +4.84% for the same model.
Visit ByteDanceResults
- Sparse MoE optimisation
- +23.3%
- Throughput, 8B fine-ranking model
- +15.36%
- Throughput from flash_attn_varlen_backward
- +4.84%
Part of Facebook Reels AI Foundation Team. Form the core AI logic of large-scale recommendation. Dabbled with ML training and experiments for different expert and foundational models. Improved Attention Kernels (fwd and bwd sdpa kernels) by > 33% resulting in training improvement > 5% qps. Implemented blockwise optimisation of SwigLU on MoE, resulting in 30% efficiency gain for SwigLU MoE, improving inference qps of 16B OneFlow model by 20% from 1490 QPS to 1790 QPS.
Visit MetaResults
- Faster SDPA kernels, fwd + bwd
- >33%
- Training QPS
- >5%
- SwiGLU MoE efficiency gain
- 30%
- Inference QPS, 16B model (+20%)
- 1490 → 1790
Built an Agentic RAG tool that automated 30% of SQL query construction for post-trade regulatory reporting. Remodelled the LLM Suite API with FastAPI, driving adoption across 5 global teams, and engineered metadata-driven vector DB switching that cut p95 retrieval latency from 6s to 2s.
Visit JPMorgan ChaseResults
- Of SQL query construction automated
- 30%
- p95 retrieval latency
- 6s → 2s
- Global teams adopting the API
- 5
Designed a PEFT fine-tuning pipeline on AWS EC2 for TinyLlama, StabilityAI, and Qwen models. Achieved a perplexity drop from 16 to 4 for Java Card code completion, with the fine-tuned model passing 90% of regex test cases versus 0% at baseline.
Visit ThalesResults
- Perplexity, Java Card completion
- 16 → 4
- Regex tests passed vs baseline
- 90% vs 0%
Re-architected a HoloLens VR app, migrating from Google Cloud to OpenAI and Anthropic and adding RAG-powered memory using gaze and pointer data for richer query relevance. Built Milvus vector endpoints that improved device energy efficiency by 20% over Pinecone, and containerised the Python server with Docker.
Visit Smart Systems InstituteResults
- Device energy efficiency vs Pinecone
- 20%
Led end-to-end development of an internal Customer Management System with Angular and Java Spring Boot, cutting decision times for 50+ sales users. Optimised contract operations by 150% via Java concurrency and delivered a data migration workflow within 2 days.
Visit CTC GlobalResults
- Contract operations optimised
- 150%
- Sales users with faster decisions
- 50+
- Data migration workflow delivered
- 2 days
Teaching & research
Conducted weekly tutorial sessions for CS2109S Introduction to AI and Machine Learning, guiding students through alpha-beta pruning, neural networks, and reinforcement learning while providing constructive feedback on assignments.
Visit NUS ComputingResearched cross-language LLM training techniques and optimised model performance through RAG and fine-tuning on GPU clusters, improving the model's ability to generate contextually relevant responses across diverse topics.
Visit NUS Computing