Engineering
175 days ago

AI / ML Engineer

Oolka · Bengaluru
This job has been flagged as a Not American
Pay
Not listed
Setting
Remote

Responsibilities:

  • Design and implement asynchronous multi-agent orchestration.
  • Own end-to-end latency from user message to AI response.
  • Build resilient inference pipelines that gracefully degrade under load.
  • Implement intelligent request routing and load balancing for AI workloads.
  • Migrate critical AI conversation flow from monolith to dedicated services.
  • Implement WebSocket/streaming infrastructure for real-time chat.
  • Design circuit breakers and fallback strategies for AI model failures.
  • Build comprehensive observability for AI system performance.
  • Optimise credit data retrieval and caching strategies.


Requirements:

  • 3-5 years building production systems handling > 10k concurrent users.
  • Proven experience with async/event-driven architectures (not just REST APIs).
  • Hands-on experience scaling ML/AI inference in production.
  • Deep understanding of caching strategies (Redis, in-memory, CDN).
  • Experience with message queues and real-time communication protocols.


AI-Specific Expertise:

  • Built systems integrating multiple LLM/AI models in production.
  • Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. )
  • Understanding of AI inference optimisation (batching, caching, model quantisation)
  • Knowledge of conversation state management and context handling.
  • Has debugged production issues under high AI inference load.
Listed by Oolka for a position based in the United States. Employers on this board attest they are hiring domestically.