⚠ SCAM ALERT: Job scams are on the rise — some use our name or logo to steal your personal or financial information. We never send links via WhatsApp, Telegram or SMS. We post only on www.jobnomix.com. Never send money to companies — jobs on JobNomix will never require payment from applicants.

Member of Technical Staff - Inference Research

Doubleword
Hybrid Full Time Other PyTorchTensorRTvLLMSGLangTensorRT-LLMAI £Not Disclosed
London, United Kingdom Hybrid Full Time Posted 1 hour ago Apply before 13 Jan 2027 28 views
Apply Now
Report this job
Thanks! Report submitted.
Please let Doubleword know you found this job on JobNomix. This helps us grow!

Stay safe — Scam Alert

Job scams are on the rise — some use our name or logo to steal your personal or financial information. We never send links via WhatsApp, Telegram or SMS. We post only on www.jobnomix.com. Never send money to companies — jobs on JobNomix will never require payment from applicants.

About the Role


We're seeking a Senior Research Engineer to join our mission of solving the hardest inference challenges in generative AI. You'll be responsible for developing cutting edge inference technology at all levels of the inference stack. This could involve writing custom kernels for inference, or designing of compute clusters for unique inference needs, or contributing to state of the art open source inference engines.

What You'll Do


Examples of projects you might work on:


  1. Building and optimizing infrastructure for high throughput inference workloads: focusing on high throughput, cost-efficient processing
  2. Inferencing fine tuned models at scale: using tools like multi LoRA and multi PEFT inference engines.
  3. Optimizing open source inference engines for offloading-based inference: implementing inference optimizations for severely memory constrained environments.

What We're Looking For


Note: A good candidate will have 80% of the following qualities. Please apply, even if the following doesn't describe you perfectly.

Core Technical Skills


  • Strong programming fundamentals
  • Understanding of GPU architectures and their performance characteristics
  • Deep understanding of LLM inference workloads, performance characteristics, and optimization techniques
  • Familiarity with Inference tooling and deep learning libraries (PyTorch, TensorRT, vLLM, SGLang, TensorRT-LLM)

Research Mindset


  • Curiosity about emerging hardware trends and ML optimization techniques
  • Ability to understand complex research requirements and translate them into infrastructure needs
  • Comfort with ambiguity and rapidly evolving technical landscapes
  • Experience supporting research workflows and experimental systems

About Doubleword


We're dedicated to making large language models faster, cheaper, and more accessible. Our infrastructure team is laser-focused on LLM inference optimization, pushing the boundaries of what's possible in terms of performance and cost efficiency while maintaining the reliability needed to serve these models at scale.
Sponsored

Match with top talent instantly

Pay $150 - $300/hour

Learn more
Sponsored

Stylish DYU Bikes On Sale

Best Seller E-Bikes Light Weight E-Bikes Enjoy Every Ride Apply Coupon Code sonakshisinghra and get 10% discount

Shop Now
Apply Now
Keep exploring

Similar Jobs

Senior Software Engineer - Zagreb

Fonoa
Zagreb, Croatia Hybrid Full Time Engineering $Not Disclosed
Hybrid Full Time Node.jsGoReact
7 min ago Apply

Associate Insurance Product Manager

Cover Genius
Malaysia Remote Full Time Insurance $Not Disclosed
Remote Full Time UnderwritersInsurersProduct Construction
2 hours ago Apply

Apply via Email

Member of Technical Staff - Inference Research · Doubleword

Send your CV and a short introduction directly to the employer.

Send email to
fergus.finn@doubleword.ai
Subject
Application: Member of Technical Staff - Inference Research — [Your Name]
Open in mail client