AI Engineer
Oolka · Bangalore, Karnataka, India
This job has been flagged as a Not American
Pay
$5k–25k
Setting
On-site
Responsibilities:
- Design and implement asynchronous multi-agent orchestration for Dhruva, our AI agent.
- Own end-to-end latency from user message to AI response.
- Build resilient inference-calling pipelines that gracefully degrade under load.
- Implement intelligent request routing and load balancing across AI workloads and providers.
- Migrate critical AI conversation flow from monolith to dedicated services.
- Implement WebSocket/streaming infrastructure for real-time chat.
- Design circuit breakers and fallback strategies for AI model failures.
- Build comprehensive observability for AI system performance, latency, error rates, and fallback triggers.
- Optimise credit data retrieval and caching strategies feeding into AI conversations.
Requirements:
- 3-5 years building production systems handling > 10k concurrent users.
- Proven experience with async/event-driven architectures (not just REST APIs).
- Hands-on experience scaling AI inference calls in production, routing, batching client-side, and retries.
- Deep understanding of caching strategies (Redis, in-memory, CDN).
- Experience with message queues and real-time communication protocols (WebSocket, SSE, Kafka or similar).
AI-Specific Expertise:
- Built systems integrating multiple LLM/AI models or providers in production, with fallback between them.
- Experience working against AI model serving endpoints (TensorFlow Serving, Triton, vLLM, or hosted LLM APIs).
- Understanding of inference optimisation from the application side: batching, response caching, prompt/context size management.
- Knowledge of conversation state management and context handling across multi-turn, multi-agent flows.
- Has debugged production issues under high AI inference load.
Listed by Oolka for a position based in the United States. Employers on this board attest they are hiring domestically.