Skip to content
Deepanshu Kr Chaurasia

Oracare AI

Clinical AI and retrieval service

Search clinical records with citations, where each patient’s data is isolated by the database itself. Retrieval is built; generation is next.

Period
Status
In development: retrieval built, generation next
Role
Software engineer and architect: service design, token validation, de-identification, retrieval, containers and Azure deployment
Team
Personal project, serving SC-Oracare
Stack
  • FastAPI
  • PostgreSQL
  • pgvector
  • Row level security
  • Microsoft Entra ID
  • Azure Container Apps
  • Azure AI Speech
  • Azure Document Intelligence
  • Docker
  • Hybrid retrieval

The problem

Clinicians using SC-Oracare dictate referrals, photograph documents and need answers from one patient’s records. Inside the web application, that work held models in memory, made slow calls to outside providers and held provider credentials the web application had no business holding. Oracare AI is the standalone service I started in August 2026 to take it over: it keeps no patient content, and it searches only within one patient’s records.

The engineering challenge

  • A different workload. Models held in memory, slow provider calls and a different scaling curve did not belong in every web server.
  • Patient content that must not travel. Nothing may be kept after a request, nothing may reach the retrieval index with an identifier on it, and no search may cross from one patient to another.
  • Proposals, not records. AI output fills a form for a person to review; it is never written as a clinical fact on its own.

Architecture

Identity provider

  • Microsoft Entra ID, Tokens, signing keys

Azure Container Apps

  • API container, FastAPI, validates every call
  • Model runtime, Holds the models

Speech and document providers

  • Speech to text, Azure AI Speech or Deepgram
  • Azure Document Intelligence, Document capture

Retrieval index

  • PostgreSQL with pgvector, Forced row level security
Validate every call
  1. Delegated token. SC-Oracare to Microsoft Entra ID (Delegated token). SC-Oracare gets a token on behalf of the signed-in clinician, scoped to one capability. A delegated token says which clinician is asking, not just which application.
  2. Full validation. Permitted. SC-Oracare to API container (Bearer token, one scope); API container to Microsoft Entra ID (Signing keys). The service checks the algorithm against an allowlist before anything else, then the signature against Entra’s published keys, found through OpenID Connect discovery, cached and refreshed when a new key appears. Then issuer, audience, expiry and the capability’s scope.
  3. Wrong scope or audience. Refused. SC-Oracare to API container (Bearer token, one scope). A token for another capability, or one minted for a different API, is refused, with the same message every time. Identity comes only from the token, never from a body, a header or a URL.
Dictation to a draft referral
  1. Transcribe. SC-Oracare to API container (Bearer token, one scope); API container to Speech to text (Audio). The audio goes to a swappable speech provider, Azure AI Speech or Deepgram, chosen by configuration.
  2. Extract. API container to Model runtime (Service token). The transcript goes to the model runtime, where a biomedical GLiNER model finds doctors, clinics, dates, allergies and medications, and rules shape them into a referral draft.
  3. Refuse, never guess. SC-Oracare to API container (Bearer token, one scope). A value that cannot be read reliably is left out and becomes a follow-up question. Nothing is kept: audio, transcript and fields live only for the request, and logs carry field names and outcomes, never values.
Capture a document
  1. Read the document. SC-Oracare to API container (Bearer token, one scope); API container to Azure Document Intelligence (Document). A photographed or uploaded document is read by Azure Document Intelligence.
  2. Propose fields. SC-Oracare to API container (Bearer token, one scope). Label-anchored rules and a confidence score for each field propose the patient’s details, which the clinician confirms before anything is filled.
Build the retrieval index
  1. De-identified export. SC-Oracare to API container (De-identified documents). SC-Oracare turns a chart into documents keyed only by record number: fields pass an allowlist, the patient’s own identifiers are removed by known value, very old ages are banded, and a final sweep aborts the whole export if any identifier survived.
  2. Write-only credential. Permitted. SC-Oracare to API container (De-identified documents). The push uses an app-only credential that can write to the index and cannot read it. The query side is the reverse, so neither credential can do the other’s job.
  3. Identifier firewall. Refused. SC-Oracare to API container (De-identified documents). Oracare AI checks every document again with a different technique, so the two checks do not share a blind spot, and refuses a document that still carries an identifier rather than cleaning it.
  4. Embed and store. API container to Model runtime (Service token); API container to PostgreSQL with pgvector (Upsert). Accepted documents are chunked, embedded in the model runtime and upserted. Pushing an unchanged document again changes nothing.
Ask about one patient
  1. Patient access grant. Permitted. SC-Oracare to API container (Bearer token, one scope). This path is built and tested in the service; the clinician chat in SC-Oracare is the next phase. Before anything is read, SC-Oracare checks its own scope and care team rules, then signs a short-lived grant bound to both the clinician and one patient.
  2. Grant mismatch. Refused. SC-Oracare to API container (Bearer token, one scope). Oracare AI verifies the grant on its own. A grant for another clinician, or for another patient than the one asked about, is refused.
  3. Search under row level security. API container to PostgreSQL with pgvector (Search one patient). The patient is set for the transaction only, and row level security, enabled and forced, is the only filter: the search query has no patient clause to forget. Vector and full-text results are fused by reciprocal rank fusion and come back with citations.
  4. Empty stays empty. API container to PostgreSQL with pgvector (Search one patient). If nothing in the chart answers the question, the result stays empty. There is no looser second attempt.
View diagram
Identity providerAzure Container AppsSpeech and document providersRetrieval indexSC-OracareServer side onlyMicrosoft Entra IDTokens, signing keysAPI containerFastAPI, validates every callModel runtimeHolds the modelsSpeech to textAzure AI Speech or DeepgramAzure Document IntelligenceDocument capturePostgreSQL with pgvectorForced row level security

Oracare AI is the third stage of the product, and the first part of SC-Oracare to become a separate service. It has its own repository, its own deployment and its own authentication, and SC-Oracare holds no model, no provider key and no provider SDK. The service does four things for SC-Oracare: transcription, medical entity extraction, document capture, and retrieval over de-identified patient records. Only SC-Oracare calls it, from its server.

Oracare AI is a FastAPI service that runs as two containers on Azure Container Apps. The API container scales out on its own and never loads a model. The model runtime holds the biomedical entity model and an open embedding model in memory, loads them once at startup, and accepts no call without a service token, so the model sits on its own scaling axis and never in every API replica’s memory.

Every capability follows one provider contract, so a provider can be swapped by configuration: Azure AI Speech or Deepgram for speech, Azure Document Intelligence for documents. Simulated providers exist for local development only, and any simulated answer is flagged as such, so a deployment can never quietly return made-up patient details as if they were real.

One rule outranks the others: the service keeps no patient content. Audio, images, transcripts and extracted fields exist only for the length of one request, and logs and audit records carry field names, sizes, durations and outcomes, never values.

Drafts, not decisions

AI output here is a proposal, never a record. When a clinician dictates a referral, the draft fills SC-Oracare’s referral wizard and stops at the review step; a person approves it, and nothing is submitted automatically.

The extraction is built to refuse rather than guess. A date without a year is not accepted as a date of birth, an unreadable time becomes a question rather than an approximation, and a doctor or clinic that does not match exactly is asked about instead of picked. Missing details become short follow-up questions, and nothing is written to the form until every question is answered, so cancelling leaves the form exactly as it was.

Retrieval and isolation

Records are isolated per patient by PostgreSQL row level security, enabled and forced, so that even the table’s owner cannot bypass it. The patient is set for one transaction only, which keeps the design safe behind a connection pooler, and the policy is the only filter: the search query has no patient clause at all, so there is nothing for the next person who writes a query to forget.

A question about a patient needs a patient access grant. SC-Oracare issues one only after its own scope and care team checks pass, keeps it short-lived, and binds it to both the clinician and that one patient. Oracare AI verifies the grant independently, so a bug on one side does not silently become a bug on both.

Search is hybrid, in one SQL statement: vector similarity with pgvector for meaning, and PostgreSQL full-text search for exact tokens such as procedure codes. The two ranked lists are fused by reciprocal rank fusion, because ranks compare cleanly where raw scores do not, and results come back with citations to the records they came from. Recency only breaks ties, so an old allergy that is still current is never buried. An empty result stays empty.

Key decisions

Delegated tokens, not app-only

Options considered: an app-only token for the whole web application, or a token SC-Oracare obtains on behalf of the signed-in clinician.

Why: an app-only token cannot say which clinician is asking, and retrieval within one patient’s records needs exactly that to decide what may be read.

Trade-off accepted: every call carries a user token the service has to validate. Filling the index is the platform acting rather than a person, so that path keeps a separate app-only credential.

Azure AI Speech in production, Deepgram for development

Options considered: Deepgram or Azure AI Speech as the production speech provider.

Why: dictated notes are patient data, so production transcription runs on Azure with the rest of the platform. The provider sits behind one contract, so the choice stays a configuration change.

Trade-off accepted: development and production use different providers, so the provider used day to day is not the one that runs in production.

A fallback in SC-Oracare for one release, then deleted

Options considered: keep the in-app models in SC-Oracare as a permanent fallback, keep them for one release, or remove them at cutover.

Why: a permanent fallback means the models never actually leave the web application. One release covers the cutover; after that, the service is the only path.

Trade-off accepted: when the service is unavailable, dictation reports an error and the form stays exactly as it was, instead of quietly falling back.

Security and identity

Each call carries a Microsoft Entra ID token that SC-Oracare gets on behalf of the signed-in clinician, scoped to one capability. The browser never talks to Oracare AI, so no token sits in page scripts, and every call still passes through SC-Oracare’s audit trail.

I wrote the validator myself. It checks the algorithm against an allowlist before it looks at the signature, then verifies the signature against Entra’s published keys, found through OpenID Connect discovery, cached, and refreshed when a token names a key it has not seen, with a cooldown so that tokens carrying random key IDs cannot make it hammer Microsoft’s key endpoint. Then it checks issuer, audience, expiry and the scope for the capability being called. Identity comes only from the token, never from a body, a header or a URL, and every rejection gets the same message.

Filling the retrieval index uses an app-only credential that can write to the index and cannot read it, while the query side can read and cannot write, so a leak of either credential is limited to one direction.

De-identification

Patient records reach the retrieval index only after two separate checks, on both sides of the boundary, using different techniques on purpose so that they do not share a blind spot.

On the SC-Oracare side, a chart becomes documents keyed only by its medical record number. Fields can reach a document only through an allowlist, the patient’s own name, contact details, address and date of birth are removed by known value, very old ages are banded as HIPAA’s Safe Harbor method requires, and a final sweep aborts the whole export if any identifier survived.

On the Oracare AI side, every incoming document is screened again, and one that still carries an identifier is refused, not cleaned. The rejection reports what kind of identifier it found, never the value.

Delivery and operations

Each capability is covered by unit tests with providers faked, so the suite runs with no keys and no network. A security suite runs across all of the capabilities together: it checks authentication and per-capability authorization, confirms that a token for one capability is refused by the others, plants a patient’s name in every capability and searches every log line for it, and confirms that forged tokens all get the same answer. A coverage floor is enforced in the test run, dependencies are audited for known vulnerabilities, and both container images run as a non-root user.

When I moved extraction out of SC-Oracare, I ran the old and new implementations side by side on the same inputs and compared them field by field before any test was written. Defects inherited from the original code were pinned by tests and reported, not fixed in the same change, so any difference in an extracted value could be traced to its cause.

Isolation is tested against a neighbouring patient given an identical vector and nearly identical text, at several retrieval depths. Retrieval quality is measured against a gold set of questions, including some that cannot be answered, with recall and abstention reported separately, and the targets the current phase did not meet are recorded rather than dropped.

Result

Both containers run in Azure, and SC-Oracare calls them for dictation and document capture. The retrieval index holds synthetic data only, and stays that way until it has a customer-managed key and a private endpoint. Retrieval with citations is built and measured; no language model is attached yet. The next phase is grounded answering on a self-hosted open model inside the same tenant, so clinical text never leaves it, followed by the clinician chat in SC-Oracare.

Demo

Oracare AI has no public endpoint: it answers only SC-Oracare, from its server. Its dictation and document capture run inside the SC-Oracare live demo (opens in new tab), and the isolation demo on the home page simulates per-patient retrieval with synthetic records.