AI engineer.
Deep in the stack, close to the customer.

JingshengChen.

Scroll ↓

Basically, I do

AI Product Engineering

AI Inference Engineering

Backend Engineering

Solution Architecture

Nice to meet you — let’s chat

01 — Work

Stuff I’ve
built.

Internships, research and open source — scroll sideways.

Solution Architect Intern

Amazon Web Services

A serverless GenAI workflow that auto-generates personalized marketing sites for exec-level clients — now used across the London AWS account-manager team.

2-week job → 5 min

  • Bedrock
  • Lambda
  • API Gateway
  • SQS
  • CDK

Backend Engineer Intern

Thought Machine

A full-stack cost tool, built from scratch, that finds idle and underused resources across the company's GCP and AWS.

flagged $18k/mo in waste

  • Go
  • Kubernetes
  • GCP
  • React
  • TypeScript

Edinburgh BSc Dissertation · Distinction

Optimising Scheduling for Batched LLM Jobs

Extended the ServerlessLLM (OSDI ’24) framework: prefetch weights over GPU DMA and model switching as a two-machine flow-shop — higher GPU utilization, lower switch overhead.

Makespan 66× faster

  • Python
  • PyTorch
  • CUDA
  • ServerlessLLM

Personal · live ↗

London BtR Map

A site for finding premium London Build-to-Rent — 300+ listings with commute routes, nearby stations and building amenities (gym / pool), plus an embedded AI property agent (tool calling + ReAct).

Real users · completion 30% → 95%

  • TypeScript
  • Next.js
  • Vercel AI SDK
  • ReAct

Open-source · via MoE-CAP

vLLM

Per-forward-pass expert-activation profiling for Qwen3-MoE — which experts fire, their counts and ratios — for easy MoE benchmarking.

Qwen3-MoE merged

  • Python
  • CUDA
  • MoE
  • Tensor Parallel
Open for 2027 internship/graduate roles* jingsheng.c04@gmail.com * Say hello *