Senior ML Engineer | Kimchi (LLM Inference Optimization)

    CAST AI

    Global RemoteNew
    ML Engineering
    AI Engineering

    About the role

    CAST AI is hiring a Senior ML Engineer to lead inference optimization for Kimchi, its system that automatically matches LLM workloads to the most cost-efficient serving configuration on customer infrastructure. The role is built around three metrics: throughput, latency, and KV cache utilization. You will tune kernels, ship quantization schemes, and adjust schedulers, with the results showing up directly in customer p99 latency and company margins. This is a high-autonomy seat where you set the technical direction rather than execute an existing roadmap.

    The work spans continuous batching, speculative decoding, chunked prefill, paged attention, prefix caching, eviction policies, quantized KV, and distributed inference topologies. You will profile TTFT and TPOT separately, identify real bottlenecks across compute, memory bandwidth, scheduling, and networking, and measure quality regressions on real workloads rather than relying on perplexity benchmarks.

    Highlights

    • Own inference optimization direction across vLLM, SGLang, and TensorRT-LLM
    • Work on KV cache, quantization, cold starts, and multi-node scaling
    • Remote-first global team with equity, learning budget, and 10% personal project time
    • Requires 5+ years building production ML systems and strong Python

    This recap is dataskew's editorial summary, not the company's copy.