MSCS (Honors) @ USC · ML / AI Engineer

Hi, I'm Aditya Jain.
I build machine learning that ships.

I'm a machine learning and software engineer on the Search team at Salesforce, where I build data and ML systems at scale — from Search Analytics, which streams tens of millions of rows per org from Apache Iceberg into customers' Data Cloud, to entity-prediction models on Salesforce's open-source ml4ir. I enjoy turning research ideas into production features people actually use.

I earned my MS in Computer Science with Honors (4.0 GPA) from USC, where I also TA'd Applied NLP (CSCI-544). Before that I spent two years as a Data Scientist at Cognizant working on search-ad click prediction and healthcare analytics. My interests span NLP, information retrieval, and computer vision — especially applications at the intersection of language and vision.

Portrait of Aditya Jain

Writing

Blog

How GPT Works — Part 5: From Base Model to ChatGPT

The final step: how a raw next-token predictor becomes a helpful assistant. Part 5 covers pretraining, supervised fine-tuning, and RLHF, the difference between a base model and a chat model, plus the context window and KV-cache that govern inference.

How GPT Works — Part 4: Training & Generation

How a transformer learns and how it writes. Part 4 covers next-token cross-entropy training with an interactive loss-descent demo, then decoding strategies — greedy, temperature, top-k, and top-p sampling — you can reshape live.

How GPT Works — Part 3: The Transformer

How the attention mechanism becomes a working language model. Part 3 covers subword tokenization with live Byte-Pair Encoding, the transformer block (residuals, LayerNorm, MLP), the causal mask that makes a GPT decoder-only, and the full pipeline from text to next-token probabilities.

How GPT Works — Part 2: Attention

From the seq2seq bottleneck to the mechanism that replaced recurrence entirely. Part 2 builds attention from the ground up — soft alignment, scaled dot-product self-attention with Q/K/V, multi-head attention, and why a transformer needs positional encoding.

How GPT Works — Part 1: The Foundations

A visual, hands-on guide to how large language models work. Part 1 covers the only prerequisites you need — vectors, the dot product, matrix multiplication, and softmax — then the one idea the whole model is built on: next-token prediction.

Toolbox

Skills

Languages

  • Python
  • Scala
  • Java
  • C / C++
  • SQL
  • JavaScript
  • HTML / CSS

ML / AI

  • Machine Learning
  • Deep Learning
  • Reinforcement Learning
  • Statistical Modelling
  • Descriptive & Inferential Statistics

Frameworks & Libraries

  • Keras / TensorFlow
  • scikit-learn
  • pandas
  • matplotlib / seaborn
  • NLTK
  • pySpark

Tools & Platforms

  • Git
  • Docker / Swarm
  • gRPC
  • MongoDB
  • Linux
  • Web Development
  • Android

Career

Experience

  1. May 2023 — Present

    MTS Software Engineer

    Salesforce, Inc.

    San Francisco, California

  2. Jan 2023 — Apr 2023

    Software Engineer

    TaxBit, Inc.

    Seattle, Washington

  3. Aug 2022 — Dec 2022

    Teaching Assistant — Applied NLP

    USC Viterbi School of Engineering

    Los Angeles, California

  4. May 2022 — Aug 2022

    Software Engineering Intern

    Salesforce, Inc.

    San Francisco, California

  5. Feb 2021 — May 2022

    Student Research Assistant

    USC Institute for Creative Technologies

    Los Angeles, California

  6. Sep 2018 — Dec 2020

    Associate Data Scientist

    Cognizant Technology Solutions

    Bengaluru, India

  7. Apr 2016 — Jul 2016

    Intern — MEAN Stack Developer

    Heelium Sports Pvt. Ltd.

    Pune, India

Academics

Education

  1. Jan 2021 — Dec 2022

    M.S. in Computer Science (Honors)

    University of Southern California

    Los Angeles, California

  2. Aug 2014 — Jun 2018

    B.E. in Computer Science

    Maharashtra Institute of Technology

    Pune, India

Say hello

Get in touch

Have an opportunity, a question, or just want to talk ML? Drop a message and I'll get back to you.