Intern - GenAI Benchmarking (MLE) / On-Device Model Deployment (SWE)

Seoul, South KoreaPosted Jul 20, 2026
## Company: Qualcomm Korea YH ## Job Area: Interns Group, Interns Group > Interim Engineering Intern - SW ## Qualcomm Overview: Qualcomm is a company of inventors that unlocked 5G ushering in an age of rapid acceleration in connectivity and new possibilities that will transform industries, create jobs, and enrich lives. But this is just the beginning. It takes inventive minds with diverse skills, backgrounds, and cultures to transform 5Gs potential into world-changing technologies and products. This is the Invention Age - and this is where you come in. General Summary: Qualcomm AI Research advances Generative AI for the edge. Our work spans model architecture restructuring, quantization, hardware-accelerated inference, deployment tooling, and the reliable evaluation of these models — enabling LLMs, vision-language and any-to-any multimodal models, and generative vision models to run on-device. This internship is open across two complementary teams, and you will join one of them depending on your interests and expertise. The **Model Efficiency & Deployment** team builds and maintains a Python library that restructures state-of-the-art open-source model architectures to run efficiently on Qualcomm hardware. It integrates PyTorch, Hugging Face Transformers, Diffusers, and AIMET, and provides a consistent set of APIs — architecture restructuring, quantization-ready graph preparation, and inference utilities. The **Evaluation & Benchmarking** team builds the benchmarking methodologies and frameworks used to evaluate these models reliably and reproducibly, addressing the still-open problem of reliability in model evaluation and benchmarking. As an intern, you'll work alongside a team of researchers and software engineers in Seoul and San Diego, with the opportunity to complete a well-defined project during the internship. Responsibilities You will contribute to one of the following two areas, depending on your interests and expertise. Model Efficiency & Deployment \- Restructure state-of-the-art open-source model architectures (LLMs, vision-language and any-to-any multimodal models, generative vision models) to run efficiently within hardware constraints. \- Apply graph-level changes to enable static-shape graph export and fixed-size KV-cache allocation, and replace operators with NPU/DSP-compatible equivalents in the model graph. \- Implement and validate efficient inference techniques such as speculative decoding, long-context handling, and KV-cache management. \- Support quantization workflows (AIMET-based post-training quantization, mixed-precision quantization, group-wise weight quantization). Evaluation & Benchmarking \- Research how to evaluate generative model outputs reliably and reproducibly where standard metrics fall short (e.g., reasoning, long-form generation), and how to separate the true impact of model changes (e.g., quantization) from generation variability. \- Improve the efficiency and reliability of evaluation to reach trustworthy conclusions at lower cost, and analyze what benchmarks actually measure to guide the choice of evaluation methods for different use cases. \- Collaborate with software and system engineers to understand requirements, plan evaluation strategies, and define key performance metrics (KPIs). Both tracks \- Validate the accuracy of quantized or compressed models against the original (unquantized) reference models, using standard task performance and quantization-error metrics. \- Read recent research papers in your area (efficient inference and quantization, or evaluation and benchmarking methodology) and implement and validate the proposed techniques in our library/framework. \- Keep your outputs up to date with new Hugging Face Transformers releases \- the Model Efficiency & Deployment side maintaining the restructured models \- the Evaluation & Benchmarking side maintaining the decoding logic built on top of them. \- Write clear, tested,...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free