AI Engineer - Agent
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
AI Engineer – Agent Systems
We (https://www.kaon.io/) build production-grade LLM-based Agents for personalized content generation and long-term interactive experiences for millions of users every day. This role owns the full Agent development lifecycle — from system design and prototyping to large-scale production deployment and continuous optimization.
Responsibilities
-
Design and implement production-ready Agent runtime frameworks, including orchestration engines, tool-use pipelines, memory systems, and context management modules.
-
Ensure system stability, scalability, and low-latency performance under high-concurrency workloads.
-
Implement asynchronous background workflows for memory consolidation and resource scheduling.
-
Design secure tool-use frameworks for interacting with external APIs and services, including sandboxing, permission management, tracing, and error handling.
-
Build end-to-end observability systems (logging, tracing, monitoring) and evaluation pipelines covering personalization quality, memory accuracy, and multi-turn consistency.
-
Establish release workflows including canary deployment, A/B testing, and production feedback loops to drive continuous improvement.
-
Optimize inference and serving performance, including caching, batching, streaming, efficient context window usage, and task scheduling.
-
Build data pipelines for collecting, cleaning, and managing Agent training data, supporting SFT, RL, and reward model iteration.
-
Collaborate closely with the algorithms team to translate research advances into scalable production systems.
Requirements
-
Strong software engineering skills with proficiency in Python; experience building large-scale distributed systems. Familiarity with Go, Rust, or C++ preferred.
-
Hands-on experience developing and deploying LLM-based Agent systems in production environments.
-
Strong system design ability with sound judgment across reliability, latency, and cost trade-offs.
-
Experience with asynchronous systems, task queues, and high-concurrency services.
-
Deep interest in Agent memory systems, personalization, and continual learning, with the ability to bridge research and engineering.
-
High ownership mindset with the ability to drive projects end-to-end from design to continuous iteration.
Nice to Have
-
Experience with vector databases, retrieval systems, and RAG architecture design.
-
Experience optimizing LLM inference and serving (e.g., vLLM, TGI, TensorRT-LLM) in high-concurrency environments.
-
Experience building Agent evaluation frameworks, simulated interaction environments, or automated benchmarking systems.
Compensation: $140,000 – $300,000 total compensation (base + equity), depending on experience and impact.
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.