CACEIS
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global
CACEIS is the asset servicing banking group of Credit Agricole dedicated to asset managers and institutional investors. Through offices across Europe, North and South America and Asia, CACES offers a broad range of services covering execution, clearing, forex, securities lending, custody, depositary, fund administration, fund distribution support, middle-office outsourcing and issuer services. CACEIS is a consolidator in the European asset servicing market and posts sustained growth in its business activities. The group holds €5.3 trillion in assets under custody and €3.4 trillion in assets under administration (figures as of 31 December 2024) By working every day in the interest of society, we are a Group committed to diversity and inclusion and place people at the heart of all our transformations. All our job offers are open to persons with disabilities.
As part of IT Innovation projects, you will join the AI Factory/Innovation team as an AI DevOps Engineer. You will be responsible for designing, implementing and maintaining the infrastructure, CI/CD pipelines and environments required to deploy and operate AI solutions in production.
What you will do:
You will work in agile mode, closely with AI developers, Data Scientists, Solution Architects and CACEIS infrastructure teams. You will play a key role in the industrialization of AI solutions and the automation of deployment processes.
As an expert in vibe coding, you use generative AI tools (GitHub Copilot, Cursor, Claude, ChatGPT) to accelerate the creation of scripts, Infrastructure‑as‑Code configurations and CI/CD pipelines, and to quickly resolve production incidents.
Design and implementation of Cloud architecture for AI solutions (scalable, secure, optimized)
Deployment and management of Infrastructure as Code (Terraform, CloudFormation)
Implementation of CI/CD pipelines for applications and AI models
Expert use of vibe coding to generate scripts, configurations and automations
Automation of deployments and rollbacks (blue/green, canary)
Configuration of monitoring, alerting and observability for AI models in production
Management of environments (dev, staging, production) and access control
Optimization of Cloud costs and performance (GPU, compute, storage)
Support to developers on tooling and DevOps best practices
Technical documentation of infrastructures and procedures
Implementation of DevSecOps practices and security compliance
Sharing of MLOps and vibe-coding best practices with the team
What you bring:
DevOps Engineer with expertise in MLOps and mastery of vibe coding, able to build and maintain robust infrastructures for AI while automating deployments and ensuring the performance of production solutions.
Technical Skills:
At least 5 years of experience in DevOps/SRE
Demonstrated experience in vibe coding using AI tools to generate scripts, configurations and diagnose incidents
Expertise in MLOps and deployment of AI models in production
Proficiency with Cloud platforms (AWS, Azure, GCP) and Cloud AI services
Expertise in Infrastructure as Code (Terraform, CloudFormation, ARM Templates)
Strong skills in containerization and orchestration (Docker, Kubernetes, Helm)
Knowledge of MLOps tools (MLflow, Kubeflow, Weights & Biases, SageMaker Pipelines)
Proficiency in scripting (Bash, Python) and automation
Experience in monitoring and observability (Prometheus, Grafana, ELK, DataDog)
Knowledge of DevSecOps security practices
Experience with secrets and configuration management (Vault, AWS Secrets Manager)
Nice to have: Experience with GPUs and optimization of AI resources
Soft Skills
Ability to use AI to quickly diagnose and resolve incidents
Pragmatism: balance between automation and delivery timelines
Rigour in designing and securing infrastructures
Innovative mindset and continuous technology watch
Autonomy and proactivity in identifying issues
Strong service orientation and support for development teams
Collaborative and pedagogical mindset
High responsiveness when dealing with production incidents
#J-18808-Ljbffr
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global
Montreal (Administrative Region) / Global