Job Title: HPC Orchestration Architect
Industry: High Performance Computing / AI Infrastructure
Location (city, state): Dallas, TX
Assignment Type: Direct hire
Pay: $180,000-$260,000 base salary, plus a potential $50,000-$100,000 bonus
Work Schedule: Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days.
Benefits: This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.
About The Company: Our client builds advanced computing infrastructure for large-scale AI, scientific research, and other demanding workloads.
Job Description: We are seeking an HPC Orchestration Architect to shape how computing resources are deployed and managed across a distributed platform. You will lead architectural decisions involving Kubernetes, virtualization, and workload placement, working closely with specialists in compute, storage, and networking.
Key Responsibilities: - Create scalable, resilient designs for orchestrating research and AI workloads across multiple locations.
- Define how containers, virtualized systems, and compute resources should be managed within the HPC platform.
- Work with other architects to ensure orchestration designs fit the broader compute, storage, and network environment.
- Turn product needs and operating constraints into clear designs that engineering teams can implement.
- Evaluate new technologies and use performance data to guide platform improvements.
- Help engineering teams apply technical standards and deliver platform changes efficiently.
- Gather requirements from customers and stakeholders to inform practical system designs.
Qualifications: - Approximately 10-12 years of relevant architecture or engineering experience preferred.
- Strong understanding of CPUs, GPUs, and other accelerators in large-scale computing environments.
- Hands-on experience building, deploying, and scaling Kubernetes clusters.
- Knowledge of virtualization, containers, orchestration, and the performance effects of workload placement.
- Experience with HPC clusters or parallel computing, including Linux optimization or performance profiling.
- Familiarity with distributed file systems and their effect on compute workloads.
- Ability to explain architectural decisions and collaborate across technical teams. A related degree is preferred, but relevant experience will be considered.
Additional Details: Relocation assistance may be tailored to the candidate, and TN visa candidates may be considered. The anticipated interview process includes an HR screen, a hiring manager meeting, a discussion of technical experience and concepts, and an onsite visit.