layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Nscale
Houston; New York; San Francisco; Seattle
Source: Nscale careers · View original posting
From Nscale's posting. “We” and “our” refer to the employer.
Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work.
Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.
Nscale is looking for a Staff AI Engineer (Specialised) to set technical direction for part of the inference platform at the core of our AI cloud: dedicated and serverless inference, bring-your-own-model deployments, and the post-training services that sit alongside them.
You’ll be the technical authority for one or more focus areas, working across the teams that build serving, post-training, and platform. Your decisions shape the latency, throughput, and cost of the tokens Nscale serves, and whether we can say with evidence that the models we run are the right ones. You’ll take on open questions, such as how KV cache should move between GPUs, nodes, and storage, or where hand-written kernels beat the open-source defaults, and your answers become the standards others build on.
We run state-of-the-art GPU systems, and much of the open-source ecosystem hasn’t caught up with them yet. This role is for people who want to close that gap hands-on while growing the engineers around them.
How We Work
Dog years.
We move quickly and compress a lot of learning into a short time.
Don’t let perfect be the enemy of good.
Ship, measure, iterate.
Be relentless.
Own the problem end to end and see it through.
One team, one mission.
Outcomes over process, and no “not my job”.
Focus Areas
We’re hiring for depth. You don’t need all of these; we want people who are exceptional in one or more
, and we’re deliberately hiring people with different focus areas:
KV cache orchestration across GPU, host memory, and storage; cross-instance cache sharing (LMCache or similar); KV-aware routing; disaggregated prefill/decode; speculative decoding; quantisation (FP8, NVFP4, INT8/4); MoE serving
writing and tuning kernels in CUDA, Triton, CUTLASS or ROCm; attention, MoE and GEMM optimisation; multi-GPU and multi-node communication on NVLink-scale systems
designing evaluation frameworks, turning customer requirements into custom benchmarks, and producing objective model comparisons we can stand behind
fine-tuning and preference optimisation as a service; RL for LLMs (PPO/GRPO-style, reward modelling, multi-turn and tool-use RL); rollout generation, weight synchronisation, and sharing infrastructure between inference and training
serving engines (vLLM, SGLang, TensorRT-LLM), model onboarding, and OpenAI-compatible APIs that other engineers and customers build on
Compensation
$220,000 — $330,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:
Here.
Nscale does not accept unsolicited candidate submissions from recruitment agencies.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: