Experience

  • Part of AML frameworks team. Focused on ML infra which is the core part of recommendation, ads and search direction. Optimised Sparse MoE by +23.3%, increasing throughput of 8B fine ranking model by +15.36%. Optimised flash_attn_varlen_backward, increasing throughput by +4.84% for the same model.

    Visit ByteDance

    Results

    Sparse MoE optimisation
    +23.3%
    Throughput, 8B fine-ranking model
    +15.36%
    Throughput from flash_attn_varlen_backward
    +4.84%
  • Part of Facebook Reels AI Foundation Team. Form the core AI logic of large-scale recommendation. Dabbled with ML training and experiments for different expert and foundational models. Improved Attention Kernels (fwd and bwd sdpa kernels) by > 33% resulting in training improvement > 5% qps. Implemented blockwise optimisation of SwigLU on MoE, resulting in 30% efficiency gain for SwigLU MoE, improving inference qps of 16B OneFlow model by 20% from 1490 QPS to 1790 QPS.

    Visit Meta

    Results

    Faster SDPA kernels, fwd + bwd
    >33%
    Training QPS
    >5%
    SwiGLU MoE efficiency gain
    30%
    Inference QPS, 16B model (+20%)
    1490 → 1790
  • Built an Agentic RAG tool that automated 30% of SQL query construction for post-trade regulatory reporting. Remodelled the LLM Suite API with FastAPI, driving adoption across 5 global teams, and engineered metadata-driven vector DB switching that cut p95 retrieval latency from 6s to 2s.

    Visit JPMorgan Chase

    Results

    Of SQL query construction automated
    30%
    p95 retrieval latency
    6s → 2s
    Global teams adopting the API
    5
  • Designed a PEFT fine-tuning pipeline on AWS EC2 for TinyLlama, StabilityAI, and Qwen models. Achieved a perplexity drop from 16 to 4 for Java Card code completion, with the fine-tuned model passing 90% of regex test cases versus 0% at baseline.

    Visit Thales

    Results

    Perplexity, Java Card completion
    16 → 4
    Regex tests passed vs baseline
    90% vs 0%
  • Re-architected a HoloLens VR app, migrating from Google Cloud to OpenAI and Anthropic and adding RAG-powered memory using gaze and pointer data for richer query relevance. Built Milvus vector endpoints that improved device energy efficiency by 20% over Pinecone, and containerised the Python server with Docker.

    Visit Smart Systems Institute

    Results

    Device energy efficiency vs Pinecone
    20%
  • Led end-to-end development of an internal Customer Management System with Angular and Java Spring Boot, cutting decision times for 50+ sales users. Optimised contract operations by 150% via Java concurrency and delivered a data migration workflow within 2 days.

    Visit CTC Global

    Results

    Contract operations optimised
    150%
    Sales users with faster decisions
    50+
    Data migration workflow delivered
    2 days

Teaching & research

  • Conducted weekly tutorial sessions for CS2109S Introduction to AI and Machine Learning, guiding students through alpha-beta pruning, neural networks, and reinforcement learning while providing constructive feedback on assignments.

    Visit NUS Computing
  • Researched cross-language LLM training techniques and optimised model performance through RAG and fine-tuning on GPU clusters, improving the model's ability to generate contextually relevant responses across diverse topics.

    Visit NUS Computing