About Haleos
We're building the operating system for how companies are founded and scaled. Haleos develops JaneOS1 and AtlasOS1 — AI that carries context across the whole journey, not a single chat. We're solving fundamental challenges in the entrepreneurial and organizational stack with production AI systems.
The Role
Architect and implement the data engine that powers our AI systems. Design robust pipelines that collect, process, and make accessible the diverse data streams flowing through our platform — enabling increasingly sophisticated insights while maintaining enterprise-grade performance and privacy.
You Will
- Design and implement scalable data pipelines that process millions of discrete data points across our platform
- Build intelligent systems for real-time data synchronization across multiple third-party integrations (CRM, financial, communication, and development tools)
- Architect vector database solutions optimized for semantic search and contextual retrieval at scale
- Develop ETL systems that transform heterogeneous business data into structured, queryable formats
- Create automated data validation and quality assurance systems to ensure training data integrity
- Build monitoring and observability tools for data pipeline health and performance optimization
- Collaborate with ML engineers to ensure training datasets meet model requirements and performance benchmarks
- Implement privacy-preserving data isolation architectures that maintain strict user namespace separation
Must Have
- Strong experience building production-grade data pipelines and ETL systems
- Proficiency with vector databases (Pinecone, Weaviate, or similar) and embedding-based retrieval systems
- Experience designing data architectures that span multiple integration points and maintain real-time synchronization
- Demonstrated ability to optimize data systems for sub-20ms query response times
- Strong understanding of database schema design, indexing strategies, and query optimization
- Experience with modern data stack tools and frameworks (Airflow, dbt, or similar)
- Proficiency in Python and SQL; experience with TypeScript is a plus
- Familiarity with API integration patterns and webhook-based data ingestion
Nice to Have
- Experience with RAG systems and LLM data architectures
- Background in building data infrastructure for SaaS platforms with complex multi-tenancy requirements
- Knowledge of data anonymization and privacy-preserving techniques
- Experience with cloud infrastructure (AWS, GCP) and containerization (Docker, Kubernetes)
- Understanding of startup operations and business intelligence systems
Benefits & Compensation
- Competitive salary and equity package
- Health, dental, and vision insurance
- 401(k) with company match
- Flexible PTO policy
- Remote-friendly work environment
- Professional development budget
- Opportunity to shape foundational technology at an early-stage company
Due to our current development phase, additional product details will be shared during the interview process.