Role ID:14
Software Engineer - DevOps and MLOps
Location:
San Francisco Bay Area, California USA
Role Description
We are looking to recruit an exceptional Software Engineer - Development and Machine Learning Operations to build and maintain the infrastructure that supports our software development, machine learning models, and AI operations.
In this role you will:
- Design, implement, and continually improve our CI/CD pipelines to facilitate seamless code integration and deployment.
- Help define and prioritize infrastructure initiatives that will help with scaling.
- Automate infrastructure orchestration and configuration management using tools such as Kubernetes, Terraform, Crossplane, Ansible, and similar.
- Monitor and optimize system performance, availability, and security.
- Configure and maintain data infrastructure appliances.
- Troubleshoot and resolve issues related to applications, infrastructure, and deployments.
- Work closely with our development and AI teams to deliver solutions that increase efficiency and stability, while optimizing for cost.
Qualifications
Must-have:
- BS or MS in software engineering, computer science, or a related field.
- Proven experience working with Kubernetes and distributed backend services.
- Deep experience managing self-hosted CI/CD systems and driving best practices around build and release workflows.
- Experience with multi-language build systems (e.g., Bazel, Bob).
- Proficiency with cloud platforms (e.g., AWS, Azure, GCP) and containerization technologies (e.g., Docker, Kubernetes).
- Experience with automation tools (e.g., Terraform, Crossplane, Ansible, GitHub Actions, Jenkins) and version control systems (e.g., Git).
- Strong programming skills in languages such as Python, Go, or Java.
- Self-starter attitude with strong ability to identify problems, prioritize them, then plan and execute working solutions.
- Enthusiasm for working in a fast paced startup environment and eagerness to support the team on a variety of topics.
Nice-to-have:
- Experience with MLOps platforms (e.g., Ray, ZenML, MLflow, Kubeflow, Databricks, Vertex AI or SageMaker).
- Knowledgeable with distributed data systems and streaming platforms such as NATS or Kafka.
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
- Understanding of machine learning frameworks (e.g., TensorFlow, PyTorch, or Scikit-learn).
- Experience with edge computing and IoT device management.
- Knowledge of security best practices and compliance standards in AI/ML environments.
- Proficiency in database management systems (e.g., PostgreSQL, MongoDB, or Cassandra).
- Experience with infrastructure-as-code tools (e.g., Terraform, Crossplane, Pulumi).
- Knowledge of GitOps practices and tools (e.g., Argo CD, Flux).