Senior Machine Learning Engineer, LLM Inference Optimization

    Nebius

    EU Remote
    ML Engineering
    AI Engineering

    About the role

    Nebius is hiring a Senior Machine Learning Engineer for its Applied AI team at Nebius Token Factory. The role owns model and endpoint optimization from model artifacts through production deployment, spanning quantization, model compression, speculative decoding, KV-cache optimization, inference engines, serving architecture, and benchmarking. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while preserving quality and reliability.

    The position involves comparing inference engines, recommending serving configurations, diagnosing regressions, and scaling dense and mixture-of-experts inference across multiple GPU nodes. You will build reproducible benchmark harnesses and collaborate with kernel and platform engineers to trace bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. The team values strong Python and PyTorch skills, hands-on LLM or.

    Highlights

    • Own LLM and VLM endpoint optimization for latency, throughput, memory, GPU utilization, and cost per token.
    • Work with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
    • Design distributed inference architectures including request routing, PDD, multi-node inference, and expert parallelism.
    • Build reproducible benchmarks measuring TTFT, TPOT, p95/p99 latency, GPU memory, reliability, and cost per token.

    This recap is dataskew's editorial summary, not the company's copy.