About Haleos
We're building the operating system for how companies are founded and scaled. Haleos develops JaneOS1 and AtlasOS1 — AI that carries context across the whole journey, not a single chat. We're solving fundamental challenges in the entrepreneurial and organizational stack with production AI systems.
The Role
Build and maintain the infrastructure that powers our AI-driven platform. Architect scalable cloud systems that handle real-time data synchronization, manage complex integration pipelines, and ensure enterprise-grade reliability as we scale.
You Will
- Design and implement cloud infrastructure for AI-powered applications with strict performance requirements (sub-20ms response times)
- Build and maintain CI/CD pipelines for rapid, reliable deployments across multiple environments
- Architect and manage vector database infrastructure optimized for semantic search and retrieval at scale
- Implement monitoring, alerting, and observability systems to ensure platform reliability and performance
- Design and maintain infrastructure for real-time data synchronization across multiple third-party integrations
- Build automated scaling solutions to handle dynamic workloads and traffic patterns
- Implement security best practices including namespace isolation, encryption, and access control
- Develop infrastructure-as-code solutions for reproducible, version-controlled deployments
- Optimize database performance, including PostgreSQL tuning and vector database optimization
- Collaborate with engineering teams to establish DevOps best practices and improve deployment velocity
- Manage disaster recovery strategies and ensure high availability across all critical systems
Must Have
- 4+ years of experience in DevOps, SRE, or infrastructure engineering
- Strong proficiency with cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code tools (Terraform, Pulumi, or similar)
- Experience with container orchestration (Kubernetes, Docker) and microservices architecture
- Solid understanding of CI/CD pipelines and automation tools (GitHub Actions, GitLab CI, Jenkins, or similar)
- Proficiency with monitoring and observability tools (Datadog, New Relic, Prometheus, Grafana, or similar)
- Experience with database management and optimization (PostgreSQL, Redis, or similar)
- Strong scripting skills in Python, Bash, or similar languages
- Understanding of networking, security, and access control best practices
- Experience with version control systems and GitOps workflows
Nice to Have
- Experience deploying and managing vector databases (Pinecone, Weaviate, Qdrant, or similar)
- Background in infrastructure for AI/ML systems or LLM applications
- Experience with real-time data pipeline infrastructure (Kafka, RabbitMQ, or similar)
- Knowledge of API gateway management and rate limiting strategies
- Familiarity with compliance frameworks (SOC 2, GDPR, HIPAA)
- Experience with multi-tenant SaaS architecture and namespace isolation
- Understanding of cost optimization strategies for cloud infrastructure
- Experience with serverless architectures and edge computing
- Background in high-availability systems with strict uptime requirements
- Knowledge of webhook infrastructure and event-driven architectures
Benefits & Compensation
- Competitive salary and equity package
- Health, dental, and vision insurance
- 401(k) with company match
- Flexible PTO policy
- Remote-friendly work environment
- Professional development budget
- Opportunity to shape infrastructure at an early-stage company
Due to our current development phase, additional product details will be shared during the interview process.