Mlops Engineer / Devops Engineer
Chefman
Mahwah
MLOps Engineer / DevOps Engineer
*Candidates must be legally authorized to work in the United States on a permanent and ongoing basis without the need for current or future employer-sponsored visa support, including H-1B, OPT, STEM OPT, or any other work authorization requiring sponsorship. Applications from candidates requiring sponsorship now or in the future will not be considered.
- Design, implement, and maintain scalable AWS cloud infrastructure supporting software, AI, and machine learning applications.
- Create and manage MLOps infrastructure for model training, deployment, monitoring, versioning, and lifecycle management.
- Partner closely with the Machine Learning Engineer to establish the tools, workflows, and infrastructure required for successful AI development and deployment.
- Support Generative AI initiatives by building infrastructure and deployment frameworks for applications utilizing AWS Bedrock, foundation models, LLMs, and related AI services.
- Build and manage Infrastructure as Code (IaC) using Terraform to ensure repeatable, secure, and scalable environments.
- Implement monitoring, logging, observability, and alerting systems across software, infrastructure, and machine learning platforms.
- Continuously identify opportunities to improve engineering processes, reduce manual effort, increase automation, and improve system reliability.
- Develop and maintain CI/CD pipelines that enable rapid, reliable software and machine learning deployments.
- Optimize cloud environments for scalability, performance, availability, and cost efficiency.
- Support security, compliance, backup, disaster recovery, and operational best practices across all environments.
- Troubleshoot infrastructure, deployment, and application issues across development, testing, and production environments.
- Document infrastructure architecture, deployment processes, operational procedures, and engineering standards.
- Contribute to establishing best practices for DevOps, MLOps, cloud architecture, and AI operations.
Please Note: Chefman is unable to provide visa sponsorship for this position. Candidates must be legally authorized to work in the United States on a permanent and ongoing basis without the need for current or future employer-sponsored visa support, including H-1B, OPT, STEM OPT, or any other work authorization requiring sponsorship. Applications from candidates requiring sponsorship now or in the future will not be considered.
- 5+ years of experience in MLOps, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or related software engineering roles.
- Strong hands-on experience with AWS services and cloud-native architecture.
- Experience building and supporting AI and machine learning platforms in AWS environments.
- Experience supporting AI, machine learning, and Generative AI applications in production environments.
- Experience working with AWS Bedrock, Generative AI services, foundation models, LLM-powered applications, or related AI infrastructure.
- Strong experience implementing Infrastructure as Code using Terraform.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Strong experience building and maintaining CI/CD pipelines and deployment automation.
- Experience supporting machine learning workflows, model deployment, monitoring, and MLOps platforms.
- Strong programming and scripting skills in Python, Bash, or similar languages.
- Experience with monitoring, logging, observability, and operational tooling.
- Strong troubleshooting, systems-thinking, and problem-solving abilities.
- Excellent communication and cross-functional collaboration skills.
- Proven track record of improving engineering processes, increasing operational efficiency, and scaling software platforms.
- Highly organized and detail-oriented with a passion for automation and continuous improvement.
- Experience with vector databases, retrieval-augmented generation (RAG), model serving, and AI infrastructure.
- Experience supporting connected devices, IoT platforms, embedded systems, or consumer technology products.
- Domain expertise in machine learning infrastructure, AI platforms, consumer applications, connected products, or similar technology environments.
- Experience working in fast-paced startup or high-growth product organizations.
Similar jobs
Ai And Data Engineer
At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life be...
Senior Mlops Engineer | $165K-$175K + Hybrid + Equity | Ai Powered Outage Intelligence Saas Startup
AdditionalinformationAdditional InformationAbout SaaS TalentSaaS Talent is more than just a recruiting company. We're your hiring, busine...
Sr. Specialist Solutions Architect -Ai&Ml Engineer
FEQ427R400Mission As a Sr Specialist Solutions Architect (SSA) - ML & AI Engineer, you will be the trusted technical ML & AI expert to bo...
Ai/Ml Engineer (Expert)
AdditionalinformationAdditional InformationWork EnvironmentNormal office conditions.Working at SOSiAll interested individuals will receiv...
Senior Manager, Forward Deployed Engineering (Uae)
Req ID: CSQ327R55Location: UAEYou will be based in the UAE and should be willing to relocate there or remain there.Databricks is redefini...
Machine Learning Engineer
At Optimove, we believe people are capable of more than a single job description. You’re not hired just to fill a position- you’re empowe...