Job description
Role Overview
AI Systems Architect - LLM & Vector Infrastructure at Star. This role designs and implements AI-native application cores where Large Language Models, vector databases, retrieval systems, and agent frameworks form the primary computational layer of web and mobile applications.
Role Purpose
Design and implement scalable AI pipelines, retrieval-augmented generation systems, memory architectures, AI agents, and orchestration workflows integrated with the development stack. The role treats AI as the operating system of the product, not as a feature.
Key Responsibilities
AI Core Architecture Design
- Design AI-first system architecture for web and mobile applications
- Architect RAG pipelines using vector databases
- Define long-term memory, short-term memory, and contextual state systems
- Implement multi-agent AI systems
- Design AI orchestration layers
Vector Database & Embedding Systems
- Select and implement vector databases including Pinecone, Weaviate, Qdrant, Milvus, and Supabase (pgvector)
- Optimize embedding strategies
- Implement hybrid search combining semantic and keyword approaches
- Design scalable indexing pipelines
LLM Integration & Optimization
- Integrate with LLM models including OpenAI APIs, Anthropic, Meta LLaMA, DeepSeek, and Alibaba Qwen
- Implement structured output pipelines
- Design evaluation and prompt testing frameworks
- Optimize cost-performance ratio
AI Agent Systems & Orchestration
- Build autonomous AI agents
- Design tool-calling systems
- Integrate with n8n and LangGraph / LangChain style agent flows
- Implement memory-aware agents
Production AI Engineering
- Build monitoring systems for hallucination detection
- Design guardrails and validation layers
- Implement evaluation datasets and benchmarking
- Ensure security of AI pipelines
- Build scalable infrastructure using Docker, Kubernetes, and GPU optimization
Qualifications & Experience
- 5+ years software engineering experience
- 2+ years building production AI systems
- Deep knowledge of vector embeddings and similarity search
- Deep knowledge of RAG architectures
- Deep knowledge of tokenization and context window optimization
- Deep knowledge of fine-tuning and LoRA concepts
- Deep knowledge of prompt evaluation frameworks
- Experience designing distributed systems
- Experience with microservices and event-driven architecture
- Experience deploying LLM systems in production
Skills & Competencies
- Python (mandatory)
- FastAPI or backend services
- Scalable API design
- PostgreSQL with pgvector
- Docker
- Kubernetes
- GPU optimization