About the Role
Our client is seeking a Senior AI/ML Engineer to build and operate production-grade Large Language Model (LLM) systems, knowledge graphs, and embedding-based retrieval pipelines for healthcare revenue cycle management. The ideal candidate will have hands-on experience deploying self-hosted LLMs, designing embedding and graph-based retrieval, and building evaluation and monitoring pipelines. This role offers full ownership of applied LLM infrastructure in a regulated healthcare environment, solving complex domain-specific problems spanning coding, claim edits, denials triage, appeal generation, and payer-rule reasoning.
Key Responsibilities
• Deploy, fine-tune, and operate self-hosted LLMs such as Llama, Qwen, MedGemma, using vLLM, SGLang, TensorRT-LLM.
• Own fine-tuning workflows (SFT, LoRA, QLoRA, DPO) on clinical notes, claims, and payer-rule data.
• Optimize GPU usage, latency, batching, and cost for production LLM inference.
• Design, maintain, and leverage knowledge graphs encoding ICD-10-CM, CPT, HCPCS, modifiers, HCC, NCCI edits, LCD/NCD policies, and payer rules.
• Build embedding-based retrieval pipelines over clinical notes, claims, denial reasons, and payer-policy corpora.
• Combine graph traversal and dense retrieval to ensure auditable, grounded outputs.
• Maintain ingestion, versioning, and quality of underlying knowledge sources (CMS, AHA, AMA, NCCI, payer bulletins).
• Build continuous evaluation pipelines with offline benchmarks and online monitoring for drift, regressions, hallucinations, and output quality.
• Track business metrics: coding accuracy, denial rate impact, clean-claim rate, cost per chart, end-to-end latency.
• Design LLM prompts, context pipelines, structured outputs (JSON, function calls, constrained decoding).
• Implement Retrieval-Augmented Generation (RAG) pipelines over medical coding standards and payer policies.
• Build MCP servers and multi-step agentic workflows with audit trails and human-in-the-loop checkpoints.
• Define deterministic vs. LLM-based tool boundaries for reliable AI-assisted workflows.
Required Skills
• 5+ years of ML/AI engineering experience.
• 6+ months of production experience with LLM systems.
• Hands-on deployment of self-hosted LLMs (vLLM, SGLang, TensorRT-LLM, or equivalent).
• Expertise in embedding-based retrieval and/or knowledge graph design.
• Experience owning evaluation infrastructure: offline benchmarks, online monitoring, drift/regression detection.
• Strong Python, PyTorch, and Hugging Face experience.
• Production experience in monitoring, incident response, and system ownership.
Nice-to-Have Skills
• Fine-tuning workflows: SFT, LoRA, QLoRA, DPO on domain-specific corpora.
• Graph databases (Neo4j, ArangoDB) and graph-aware retrieval.
• Vector databases, hybrid search (BM25 + dense, rerankers).
• Familiarity with LLM observability tools: Langfuse, LangSmith, Arize, Braintrust, or in-house equivalents.
• Experience in healthcare, RCM, claims, or regulated domains.
• Experience with MCP or similar tool orchestration frameworks.
• Strong prompt-engineering and LLM evaluation skills.
About YMinds.AI
YMinds.AI is a technology and AI talent partner, helping organizations hire exceptional professionals across AI/ML, Data Science, Cloud Engineering, Full Stack Development, Product Engineering, and emerging technology domains. Using the proprietary EmployAbility.AI platform, YMinds.AI delivers pre-vetted, highly skilled talent to accelerate innovation and scale teams efficiently.
Keywords
Senior AI Engineer, ML Engineer, LLM Engineer, Large Language Models, Self-Hosted LLM, vLLM, SGLang, TensorRT-LLM, PyTorch, Hugging Face, Knowledge Graph, Embedding-Based Retrieval, RAG, Healthcare AI, Revenue Cycle Management, MCP, Prompt Engineering, Denial Management, Claim Automation, Medical Coding, ICD-10, CPT, HCPCS, NCCI
Hashtags
#AIEngineer #MLEngineer #LLM #HealthcareAI #RevenueCycleManagement #SelfHostedLLM #KnowledgeGraph #PromptEngineering #RAG #PyTorch #HuggingFace #MedicalCoding #DenialManagement #AIJobs #TechHiring #YMindsAI

