About This Role
Gen AI Inferencing Engineer
Full-Time | On-Site | Dallas, TX
---
About the Company
A leading generative AI inferencing company is at the forefront of deploying large-scale AI systems that power real-world, production-grade applications. With a deep commitment to infrastructure excellence and platform reliability, this organization builds the foundational tooling that enables data science and machine learning teams to move fast and ship with confidence. This is a place where infrastructure is treated as a first-class product.
---
The Role
This is a senior infrastructure and MLOps engineering role built for someone who has spent meaningful time running machine learning models in production — at scale, under pressure, and with real performance constraints. You will own the inferencing platform that internal teams depend on, driving decisions around deployment architecture, optimization, and reliability. If you come from an ML platform, MLOps, or SRE-for-ML background and have hands-on experience serving generative AI workloads specifically, this role was designed with you in mind.
---
What You Will Do
• Design, deploy, and maintain high-performance generative AI inferencing infrastructure capable of handling demanding production workloads
• Own containerized model serving environments built on Docker and Kubernetes, ensuring stability, scalability, and repeatability across deployments
• Tune inference systems for throughput and latency, identifying bottlenecks and implementing targeted optimizations that move the needle on real performance metrics
• Build and maintain MLOps CI/CD pipelines that support model fine-tuning workflows and smooth promotion of models from development into production
• Develop and steward shared ML platform tooling that multiple data science and engineering teams build on top of — not one-off solutions, but durable infrastructure
• Collaborate closely with research and applied science teams to translate model requirements into reliable, scalable serving configurations
• Contribute to the evolution of inference framework internals, staying current with the rapidly shifting generative AI infrastructure landscape
---
What We Are Looking For
• Proven experience with vLLM or Triton Inference Server in production
• Strong containerization skills across Docker and Kubernetes
• Hands-on MLOps background: CI/CD pipelines for ML workflows
• Deep experience with throughput and latency tuning for inference systems
• Familiarity with inference framework internals and how they affect serving behavior
• Experience supporting fine-tuning workflows end to end
• Track record building shared ML platform tooling used across teams
The ideal candidate is an infrastructure-first engineer whose primary domain is making models run reliably and efficiently at scale — not building applications on top of them. You have likely worked in an ML platform, MLOps, or reliability engineering context where generative AI serving was a core part of your mandate. You understand the difference between getting a model to run and getting it to run well in production, and you have the scars to prove it. Experience with traditional ML model serving alone is not sufficient — what matters here is direct exposure to the unique demands of large language model and generative AI inferencing.
---
Nice to Have
• Experience implementing Retrieval-Augmented Generation (RAG) pipelines
• Ability to design and evaluate retrieval logic within AI serving architectures
Familiarity with RAG is a genuine advantage in this role, as the platform increasingly supports use cases where retrieval and generation are tightly coupled. Candidates who have thought carefully about how retrieval logic integrates with inference infrastructure — not just at the application layer, but in terms of latency budgets and system design — will find immediate opportunities to apply that knowledge. This is secondary to core inferencing expertise but meaningfully differentiates strong candidates.
---
What We Offer
This is an on-site role based in Dallas, TX, embedded within a team that treats infrastructure as a strategic asset rather than a support function. You will have real ownership over systems that matter, with the autonomy to make architectural decisions and the visibility that comes from building foundational tooling others depend on. The generative AI inferencing space is evolving rapidly, and this position offers direct exposure to cutting-edge serving technologies and the opportunity to grow alongside a field that is redefining what production AI looks like.