1
Kategoria
Skadon
Orari
Lokacioni
	 	 

Applied AI Engineer

Omzo is hiring a Staff Applied AI Engineer to build the AI backend for Omzo Air, from managed-model evaluation through a safe, measurable production launch for a health platform operating under HIPAA requirements.

About Omzo

Omzo is building Omzo Air, an AI assistant that helps users understand approved health and wearable information, navigate appropriate next steps, and connect with a human care team when necessary.

Omzo Air is not intended to operate as an autonomous clinician. We are building for production from the beginning, with measurable safety, reliable performance, controlled costs and clear clinical boundaries.

About the role

We plan to use a managed model because it is the most practical way to meet our HIPAA, security and reliability requirements without operating our own LLM infrastructure.

The OpenAI API and Azure OpenAI are our current leading candidates, but the final deployment has not been selected. You will lead the technical evaluation and recommend the provider, model and configuration based on quality, safety, latency, regional availability, cost and operational fit.

Privacy, Security and Legal will approve the contractual and data-handling requirements. Clinical leadership will approve the intended use, clinical-safety rules and release criteria.

You will build and own the AI backend service. Omzo’s backend, mobile and web engineers will call the internal API you provide and own the surrounding application and user-interface code. Platform/SRE owns cloud infrastructure, networking, deployments and infrastructure incidents.

Healthcare experience is helpful but not required. Experience working with domain experts on a regulated, safety-critical or sensitive-data system is important. You will not be expected to make clinical-policy decisions independently.

What you will build and own

  • Lead the technical evaluation of OpenAI API, Azure OpenAI and any approved alternatives.
  • Build the AI backend service and a stable internal API for Omzo’s backend, mobile and web teams.
  • Build the model gateway, including routing, streaming, timeouts, bounded retries, rate limits, approved caching, cost tracking and safe degraded behavior.
  • Own model-related response performance, working with Platform/SRE on infrastructure capacity, networking and provider quotas.
  • Build retrieval over clinician-approved sources, including ingestion, provenance, review status, versioning, citations and rollback.
  • Implement task and risk routing that can:
  • Use a smaller, faster model for validated tasks such as intent classification, retrieval preparation and approved FAQs.
  • Use a more capable model for complex tasks that remain within the approved product scope.
  • Refuse or limit requests that fall outside the approved scope.
  • Escalate high-risk, uncertain or clinician-only questions to a human care team.
  • Design for cost efficiency at every layer, and treat the model as the last resort rather than the default:
  • Handle predictable requests with plain application code instead of a model call: greetings, menu-style navigation, appointment and shipping status, account questions, and any answer that can be looked up directly from our systems.
  • Serve approved FAQs and repeated questions from a reviewed answer library or a semantic cache, with the model used only when no approved match exists.
  • Use deterministic rules and lightweight classifiers before any model call, so simple intents never reach a paid endpoint.
  • Keep prompts and context small: cached system prompts, bounded conversation history, only the retrieved sources the answer needs, and capped output length.
  • Track cost per successful task by route, and regularly move workloads from the model to code, from larger to smaller models, or from live calls to batch processing when evaluations show no loss in quality or safety.
  • Build layered safety controls using clinician-approved rules, classifiers and conversation context.
  • Build response-release controls. Policy-sensitive or high-risk responses must be checked before release, using buffered or validated segment streaming where appropriate.
  • Build the human-escalation workflow, including ticket creation, priority, minimized context, delivery confirmation, retries, status tracking and an approved interim response for the patient.
  • Build clinician-reviewed evaluation suites covering retrieval, grounding, citations, risk detection, routing, refusal behavior, supported languages, prompt injection and regressions.
  • Run evaluations for every relevant change and on a schedule using synthetic, de-identified or otherwise approved cases.
  • Build privacy-safe monitoring for model quality, latency, reliability, escalation behavior and cost without exposing protected health information in ordinary logs or traces.
  • Lead model and prompt versioning, shadow evaluations, canary releases and rollback in partnership with Platform/SRE.
  • Review production quality weekly with the clinical lead and turn findings into measurable engineering improvements.
  • Establish standards for the AI codebase and mentor engineers who join the AI work later.

A stronger model is not a substitute for a doctor. Routing must consider clinical risk separately from technical complexity, and every route must pass its own safety and quality evaluations. Cost savings never override clinical-safety thresholds.

Fallback traffic may only be sent to model deployments that have independently passed contractual, regional, security and clinical-evaluation requirements. If no approved fallback is available, the system must return a safe degraded response and provide the appropriate human or emergency path.

What we are looking for

  • Demonstrated Staff-level technical leadership and deep hands-on ownership of at least one production LLM system used by real customers, typically supported by six or more years of relevant engineering experience.
  • Production experience with structured outputs, retrieval-augmented generation, model evaluation and managed model APIs.
  • Experience building internal APIs or backend services used by other application teams.
  • Experience designing reliable services around databases, queues, rate limits, timeouts, retries and external API dependencies.
  • A track record of using evaluations to make model, prompt and routing decisions.
  • Experience reducing AI spend in production through routing, caching, prompt design or replacing model calls with code, without lowering quality.
  • Experience implementing classifiers, validation, refusal, escalation and conservative failure behavior.
  • Experience protecting sensitive information in a regulated or security-conscious environment.
  • Ability to lead technical design and explain trade-offs clearly to engineering, product, clinical and leadership stakeholders.

How we work

  • We release in phases with clear entry and exit criteria.
  • Every model, prompt, route, retrieval, policy or knowledge change must pass its evaluation gates before release.
  • We prefer measured evidence over provider benchmarks or intuition.
  • We use clinician-reviewed scoring rubrics rather than assuming every health question has one exact answer.
  • We review production quality with the clinical lead every week.
  • We document important provider, architecture and clinical-safety decisions.

Omzo is an equal-opportunity employer. We welcome candidates from different backgrounds and provide reasonable accommodations throughout the interview process.

How to apply

Send your CV in English to indrit@omzo.com and a short example of a production LLM system you built. Explain the problem, how you designed the routing, retrieval and safety controls, how you evaluated it before launch, and one thing that went wrong in production and what you changed. Please use public, anonymized or sample details rather than confidential information.

KosovaJob është rrjeti më i madh i punësimit në Kosovë i çertifikuar nga Bureau Veritas me ISO 9001:2015 Standardet për kualitet