Position Overview
We are hiring an AI Architecture & Infrastructure Engineer to design and build bank-grade AI infrastructure, model platforms, and intelligent-application foundations. The role is responsible for turning AI capabilities into secure, reliable, scalable, auditable, and operationally sustainable enterprise platforms, covering data processing, model training, fine-tuning, evaluation, deployment, online inference, monitoring, and lifecycle management.
The position focuses on AI architecture, distributed computing, GPU resource management, Kubernetes, MLOps/LLMOps, model serving, cloud-native engineering, security hardening, reliability, and business continuity. In this job description, “hardness engineering” is interpreted as security hardening, reliability engineering, and resilience engineering for AI infrastructure.
Item
Details
Job title
AI Architecture & Infrastructure Engineer
Function
AI Platform, Technology Architecture, Cloud Infrastructure, and Machine Learning Engineering
Reporting to
Head of AI Platform / Head of Technology Architecture
Location
[City / Working arrangement]
Employment type
Full-time
Number of openings
[Number]
Key Responsibilities
1. AI Architecture and Platform Engineering
Design and deliver the architecture of the bank’s AI infrastructure and platform, covering data, training, fine-tuning, evaluation, model registry, model serving, inference gateway, identity, auditability, observability, and operations. Develop architectures for offline training, batch prediction, real-time inference, high-concurrency serving, and multi-model collaboration. Establish architecture standards, interface specifications, technical baselines, and platform evolution roadmaps.
2. Model Training and Data Processing Infrastructure
Build repeatable, scalable, and auditable model-training infrastructure for supervised learning, deep learning, large language models, and other machine-learning workloads. Own training orchestration, dataset management, data preparation, feature processing, experiment tracking, model evaluation, resource scheduling, training logs, metric collection, and model-artifact management. Optimize data, compute, network, storage, and communication efficiency for training workloads.
3. Model Fine-Tuning and Large-Model Engineering
Design and implement infrastructure for fine-tuning large language models and domain-specific models, supporting instruction tuning, supervised fine-tuning, parameter-efficient fine-tuning, and relevant preference-optimization workflows. Understand and apply techniques such as LoRA, QLoRA, adapters, quantization, distillation, mixed precision, gradient accumulation, checkpoint management, and distributed training. Integrate these capabilities into standardized training, evaluation, release, and rollback processes. Work with data, model-risk, business, and security teams to ensure that training data and model artifacts meet internal governance requirements.
4. Kubernetes and Cloud-Native Platform Engineering
Build and operate AI workloads on Kubernetes (K8s), including containerization, Pod scheduling, GPU allocation, node-pool management, autoscaling, job queues, network policies, service discovery, storage orchestration, namespace isolation, quota management, and multi-tenant governance. Work with Helm, Operators, Ingress, service mesh, CI/CD, GitOps, Terraform, or equivalent technologies. Troubleshoot reliability, performance, scheduling, and resource-isolation issues for training and inference workloads running on Kubernetes.
5. Model Serving and Inference Optimization
Build a unified model-serving platform and model gateway for conventional machine-learning models, deep-learning models, and large language models. Own model containerization, version management, canary releases, A/B testing, rate limiting, circuit breaking, caching, batch inference, asynchronous invocation, and multi-model routing. Continuously optimize inference latency, throughput, GPU/CPU utilization, memory consumption, concurrency, and cost per request. Experience with model quantization, continuous batching, KV cache, inference caching, or other acceleration techniques is preferred.
6. MLOps/LLMOps and Delivery Automation
Build automated workflows covering data preparation, training, fine-tuning, evaluation, model registration, approval, release, and production monitoring. Maintain traceability among model versions, dataset versions, code versions, configuration parameters, and experiment results to ensure reproducibility, rollback, and auditability. Provide standardized SDKs, APIs, templates, pipelines, and self-service tools for data-science, machine-learning, and application teams.
7. Security Hardening, Reliability, and Resilience Engineering
Establish security baselines and production-reliability mechanisms for the AI platform, including authentication, fine-grained authorization, secrets and key management, network segmentation, data masking, sensitive-data protection, image and dependency scanning, software supply-chain security, runtime protection, vulnerability remediation, audit trails, and anomaly detection. Define service-level objectives, capacity-management practices, incident-response procedures, backup and recovery plans, business-continuity controls, disaster-recovery failover, and regular resilience testing. Improve recoverability under traffic surges, dependency failures, hardware failures, model anomalies, and security incidents.
8. Observability and Production Operations
Build unified observability across infrastructure, training jobs, model services, data pipelines, and business indicators. Use logs, metrics, distributed tracing, model-quality indicators, data-drift detection, performance metrics, and cost metrics to diagnose platform and model issues. Participate in production on-call rotations, major-incident reviews, capacity planning, performance testing, and platform upgrades, driving root-cause remediation rather than temporary fixes.
9. Technical Leadership and Cross-Functional Delivery
Own technical design reviews, proof-of-concept validation, solution decomposition, production acceptance, and documentation for key initiatives. Mentor software, machine-learning, platform, SRE, data, security, and risk engineers. Participate in technical hiring, interviews, code reviews, and architecture reviews, helping the organization build standardized, reusable, and continuously evolving AI engineering capabilities.
Minimum Qualifications
Capability area
Requirements
Education
Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, Network Engineering, Information Security, or a related field.
Experience
5+ years of experience in AI platforms, cloud platforms, distributed systems, MLOps/LLMOps, platform engineering, or related infrastructure development. Experience in banking, financial services, or other highly regulated industries is preferred.
Programming
Strong proficiency in at least one of Python, Go, Java, or C++, with solid knowledge of data structures, algorithms, concurrency, networking, and system design.
Linux and cloud-native engineering
Hands-on experience with Linux, Docker, Kubernetes/K8s, container networking, service discovery, CI/CD, GitOps, and infrastructure as code.
Kubernetes expertise
Understanding of K8s scheduling, Deployment, StatefulSet, Job, CronJob, Service, Ingress, ConfigMap, Secret, RBAC, NetworkPolicy, resource quotas, and autoscaling. Experience with GPU workloads and multi-tenant clusters is preferred.
Model training
Understanding of data preparation, training orchestration, distributed training, mixed precision, checkpointing, experiment tracking, model evaluation, and model registration.
Model fine-tuning
Experience with supervised fine-tuning, instruction tuning, LoRA/QLoRA, adapters, quantization, distillation, parameter-efficient fine-tuning, or related large-model training techniques.
AI frameworks
Experience with at least one of PyTorch, TensorFlow, JAX, Hugging Face Transformers, DeepSpeed, FSDP, or equivalent frameworks and tools.
Model serving
Understanding of model deployment, online inference, model gateways, version control, canary releases, continuous/dynamic batching, caching, rate limiting, circuit breaking, and inference-performance optimization.
Data and storage
Experience with one or more of object storage, relational databases, NoSQL, message queues, data lakes, or feature stores, together with an understanding of data access, lineage, and version governance.
Security and reliability
Experience with IAM, secrets management, network segmentation, vulnerability management, supply-chain security, auditability, disaster recovery, incident response, and security hardening.
Engineering discipline
Strong focus on automation, testing, observability, documentation, code quality, change management, and production operations; able to own delivery from design through production.
Communication
Able to communicate effectively with business, technology, data, risk, compliance, audit, and information-security stakeholders, translating complex technical issues into clear solutions and decisions.
Preferred Qualifications
Experience with core banking systems, financial data platforms, risk management, anti-money laundering, customer service, intelligent operations, or other financial AI use cases. Experience with GPU clusters, NVIDIA technologies, distributed training, inference acceleration, model gateways, retrieval-augmented generation, AI agents, model security, or AI developer platforms. Familiarity with Terraform, Helm, Argo CD, Prometheus, Grafana, OpenTelemetry, Kafka, Redis, PostgreSQL, object storage, or major cloud services. Demonstrated experience building an AI platform from the ground up, scaling platform adoption, or handling major production incidents. Relevant certifications in cloud, Kubernetes, information security, data, or related technologies are a plus.
Expected Outcomes and Success Measures
Area
Expected outcomes
Platform delivery
Establish a unified platform for AI training, fine-tuning, evaluation, deployment, and inference, reducing duplicated implementation.
Training efficiency
Improve reproducibility, resource utilization, experiment management, and delivery speed for training and fine-tuning workloads.
Inference performance
Optimize model-serving latency, throughput, memory utilization, concurrency, and cost per request.
Production reliability
Improve availability and recoverability through automation, observability, capacity management, and resilience testing.
Security and governance
Establish security baselines, access controls, audit records, and a closed-loop vulnerability-remediation process.
Developer experience
Provide standardized SDKs, templates, APIs, pipelines, and documentation to shorten the path from model development to production.
What We Offer
Join our banking AI infrastructure team and contribute to enterprise-grade AI platforms designed for production-scale adoption. You will collaborate with AI, cloud, data, security, risk, and business teams while working on highly reliable, secure, and performant systems. The role provides opportunities to influence critical technical decisions, platform evolution, and engineering capability development. Compensation, benefits, training, and career development are subject to the bank’s internal policies and applicable local regulations.
Recruitment Process
The recommended process includes resume screening, an initial technical interview, a system-design and AI-infrastructure interview, a model-training and fine-tuning interview, a Kubernetes/cloud-native technical exercise or discussion, cross-functional interviews, risk and compliance discussions, a management interview, and background checks. The bank may adjust the process based on seniority, location, and internal recruitment policies.
Application Materials
Please submit an English resume describing AI platforms, model-training, model-fine-tuning, Kubernetes, distributed-systems, cloud-native infrastructure, or security-hardening projects you have led or contributed to. Where possible, include project scale, technology stack, personal responsibilities, training or inference metrics, resource utilization, availability targets, cost-optimization results, and incident-management experience. Please anonymize confidential information relating to banks, customers, or internal systems.
Suggested Level Naming
Level
Suggested title
Mid-level
AI Infrastructure Engineer
Senior
Senior AI Architecture & Infrastructure Engineer
Principal
AI Platform Architect / AI Infrastructure Principal
Technical leadership
Head of AI Infrastructure / AI Infrastructure Engineering Lead
Before publishing: Complete the seniority level, compensation range, location, reporting line, on-call requirements, primary cloud environment, main models and frameworks, and applicable local compliance requirements.
Document author: Manus AI
References
This job description was prepared based on the hiring requirements provided and does not cite external sources.
| 薪酬 | 薪金面議 |
| 工種 |
|
| 工作地點 |
|
| 僱用形式 |
|
| 教育程度 |
|
刊登於 6日前
刊登於 5日前