Senior/Staff/Senior Staff AI Infrastructure & Applied AI Engineer
Design, optimize, and deploy large-scale AI systems including LLM inference platforms and agent-based automation architectures. Build and scale serving frameworks, optimize model performance using PyTorch and CUDA, and integrate models into production workflows. Requires strong Python skills, deep understanding of transformer architectures, and experience in distributed systems or cloud infrastructure.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
The success of any AI initiative depends on one critical factor: people.
Not just skills, but judgement, experience, and the ability to operate in real-world environments. Nexus People is a specialist consulting firm dedicated to supporting organisations build their future with top AI talent & capabilities.
General SummaryThis is an open call to identify candidates across a family of related engineering roles covering the design, optimisation, and deployment of large-scale AI systems. Roles are available at Senior, Staff, and Senior Staff level, spanning two broad areas of focus:
- Machine learning infrastructure and performance engineering, covering how large models are served, scaled, and made efficient in production.
- Applied AI engineering for agent-based and automation-driven systems, covering how models are integrated into tools, workflows, and decision-making pipelines.
Candidates working in either of these areas, or in adjacent disciplines within applied AI and machine learning systems engineering, are invited to express interest.
Key Responsibilities- Build and scale inference platforms and serving frameworks for large language, vision-language, and diffusion models.
- Contribute to LLM serving and orchestration packages such as vLLM, SGLang, TGI, Triton Inference Server, Dynamo, or LLM-d.
- Transform and optimise models for efficient inference using PyTorch, ONNX, and graph capture/compilation approaches such as torch.compile and TorchDynamo.
- Design and implement specialised kernels and fusion operations, for example in Triton or CUDA, for performance-critical workloads.
- Analyse and improve throughput and latency across batching, parallelism, KV-cache management, and speculative decoding strategies.
- Design agent architectures that decompose complex goals into executable steps, integrate LLMs with external tools, APIs, and vector databases, and implement memory, guardrail, and monitoring systems for autonomous agents.
- Profile, debug, and resolve complex performance or stability issues to root cause.
- Collaborate with compiler, firmware, platform, and customer-facing teams to move solutions from research through to commercial deployment.
- Engage with open-source AI/ML communities to evolve serving and optimisation frameworks.
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
- Strong Python development skills for large-scale software projects.
- Hands-on experience with PyTorch, and, depending on specialisation, ONNX, LangChain, or comparable frameworks.
- Deep understanding of transformer-based architectures, attention mechanisms, and mixture-of-experts models.
- Strong computer science fundamentals, including algorithms, data structures, and parallel or distributed programming.
- Understanding of computer architecture, ML accelerators, and distributed systems, or, for agentic-focused candidates, cloud infrastructure (AWS, Azure, GCP).
- Experience analysing, profiling, and optimising deep learning workloads for throughput and latency.
- Experience with RESTful APIs, vector databases, and semantic search (for agentic/applied AI candidates).
- Strong communication and problem-solving skills, with the ability to operate across the full lifecycle from research and prototyping through to commercial deployment.
- MSc in Computer Science, Machine Learning, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
- Open-source contributions to any GenAI, LLM serving, or agent-orchestration package.
- Experience architecting and developing large-scale distributed systems.
- High-level kernel design experience (PyTorch, CUDA, Triton).
- Experience with continuous batching, disaggregated serving, or KV-cache management at scale.
- Background in numerical methods, accuracy evaluation frameworks, or ML compilers.
- PhD in Computer Science, Computer Engineering, Machine Learning, or a related field.
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field and 4+ years of Software Engineering or related work experience; OR
- Master's degree in Engineering, Information Systems, Computer Science, or a related field and 3+ years of Software Engineering or related work experience; OR
- PhD in Engineering, Information Systems, Computer Science, or a related field and 2+ years of Software Engineering or related work experience.
- 2+ years of work experience with a programming language such as C, C++, Java, or Python.
- Salary, stock, and performance-related bonus
- Employee stock purchase scheme
- Matching pension scheme
- Maternity/Paternity Leave
- Education Assistance
- Relocation and immigration support (if needed)
- Life, Medical, Income, and Travel Insurance
- Subsidised memberships for physical and mental well-being
- Bicycle purchase scheme
Similar Jobs
Explore other opportunities that match your interests