All permanent roles
AI Engineer
Build the AI systems behind IICL’s own products — iVaak AI for multi-tenant voice, TruFix AI for an audit-first service desk, and iWac.AI for chat and WhatsApp Business — as well as custom work for clients. The role covers agentic workflows, retrieval pipelines, voice agents and the APIs and infrastructure that run them, owned end to end from prototype to a deployed service with its evaluation, tracing and cost under control. It suits an engineer who has already put an LLM system in front of real users and knows what breaks.
What you will do
Design and build LLM-powered applications — agentic workflows, multi-agent systems, RAG pipelines and conversational agents in voice and chat.
Develop and optimise retrieval-augmented generation: chunking strategy, embeddings, vector search, re-ranking and grounding for accuracy.
Build and orchestrate agents with LangGraph, CrewAI, LangChain or LlamaIndex, including tool and function calling and structured outputs.
Integrate and, where it is warranted, fine-tune foundation models — Claude, GPT, Gemini and open-source LLMs — choosing per use case on quality, latency and cost.
Build production backend services and APIs in Python and FastAPI that expose AI capabilities reliably at scale.
Work on voice AI features: speech-to-text, text-to-speech and real-time conversation handling across providers such as ElevenLabs, Deepgram and Twilio telephony.
Implement evaluation, observability and guardrails — prompt versioning, tracing, automated evals and hallucination and safety checks through Langfuse, RAGAS or equivalent.
Optimise latency, token cost and reliability against real production traffic rather than benchmark conditions.
Deploy through Docker and Kubernetes on AWS or Azure with CI/CD, working alongside product.
Track a fast-moving field and bring what genuinely works into the products.
What we are looking for
Three to five years in software engineering, including two or more building AI, ML or LLM-powered systems that reached production.
Strong Python and sound engineering fundamentals — clean code, tests and version control.
Hands-on experience with LLM APIs: Anthropic Claude, OpenAI, Azure OpenAI or Google Gemini.
Practical RAG pipeline experience and work with a vector database such as pgvector, Pinecone, Weaviate or Qdrant.
At least one agent or orchestration framework in anger: LangGraph, LangChain, CrewAI or LlamaIndex.
Prompt engineering, with a working understanding of context management, function calling and structured outputs.
Building and consuming REST APIs with FastAPI, Flask or similar.
Familiarity with AWS or Azure and with containerisation through Docker; Kubernetes is a plus.
A grasp of embeddings, tokenisation, model evaluation and the cost and latency trade-offs between models.
Clear communication, and the ability to own a feature from prototype through to production.
Preferred background
Voice AI experience — TTS and STT, real-time streaming or telephony through ElevenLabs, Deepgram, Twilio or Plivo.
LLM observability and evaluation tooling: Langfuse, RAGAS or LangSmith.
Fine-tuning and model customisation with LoRA or PEFT, or serving open-source models such as Llama, Mistral or Qwen through vLLM or Ollama.
Model Context Protocol (MCP), or building tool integrations for agents.
Multi-tenant SaaS architecture and data isolation.
WhatsApp Business API, CRM or ITSM integrations, or workflow automation.
Classical ML and NLP, and the data pipelines behind them.
Open-source contributions or a portfolio of shipped AI work.
