Skip to content
SWE JOB LISTS

Software Engineer Project Intern (Model Infrastructure) - 2026 Start (BS/MS) at TikTok

San Jose, CA, United StatesInternshipPosted Oct 3, 2026

Description

Seniority: Intern

TikTok is a short-form mobile video platform seeking a Software Engineer Project Intern for its Model Infrastructure team. The intern will optimize large-scale recommendation model training and inference, develop LLM-integrated recommendation infrastructure, and work on distributed streaming, hardware-aware co-design, and model state management.

Responsibilities

  • Engineering Efficiency at Scale: Drive the optimization of training and inference pipelines to maximize hardware utilization (MFU/HFU) for models featuring hundreds of billions of dense parameters
  • LLM2Rec Infrastructure: Architect specialized systems to support the integration of LLMs into the recommendation stack, focusing on memory-efficient attention mechanisms and advanced KV cache management for long-sequence user modeling
  • Massive Sparse & Dense Streaming: Build and optimize high-concurrency engines for Petabyte-scale streaming training, handling continuous parameter updates and high-frequency data ingestion without compromising stability
  • Hardware-Aware Co-Design: Work closely with researchers to design next-generation recommendation architectures optimized for modern GPU/NPU interconnects, ensuring high-bandwidth utilization across the cluster
  • Distributed State Management: Innovate on how we store and synchronize massive model states across heterogeneous memory hierarchies (HBM, DDR, and NVMe)

Qualifications

  • Currently pursuing an Undergraduate/Master in Software Development, Computer Science, Computer Engineering, or a related technical discipline
  • Strong programming skills in C++ and Python
  • Solid understanding of Computer Architecture and the GPU software stack (CUDA, Triton, or NCCL)
  • Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and a desire to "look under the hood" of model execution runtimes
  • A strong interest in solving system-level bottlenecks in large-scale distributed environments

Preferred Qualifications

  • Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP)
  • Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering)
  • Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX)
  • Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures
  • Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects

Skills

  • Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP)
  • Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering)
  • Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX)
  • Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures
  • Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects

Tech stack

C++, Python, Computer Architecture, CUDA, Triton, NCCL, PyTorch, TensorFlow, Transformer Architectures, 3D Parallelism, torch.compile, High-Performance Kernel Development, UCX, RDMA Networking

Benefits

  • Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
  • Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year).
  • Interns who are not working 100% remote may also be eligible for housing allowance.

More from TikTok

All TikTok roles →

More in San Jose

All jobs in San Jose →

Similar roles

All internships →