Dataveoli

Clinical data abstraction that never leaves the Kingdom

Saudi hospitals are legally required to report to national health registries. The abstraction behind those submissions is still done by hand, one patient chart at a time.

How it works

01

The obligation

A statutory reporting duty met with a manual process.

Cancer notification is mandatory by ministerial decree. More than ninety communicable conditions are notifiable, some immediately and others within 72 hours. Quality indicator returns and national specialty registries sit alongside them.

None of these obligations is optional, and none of them is currently automated. They differ in destination, deadline and data dictionary, and they converge on one production step: a person reads a patient record and transcribes defined variables into a standard form.

The source material is not a database. It is free-text progress notes, operative and radiology reports, scanned outside records, and mixed Arabic–English shorthand. A database query cannot produce these variables. A human reading the note can, slowly.

516

Hospitals in the Kingdom, each carrying these duties

90+

Notifiable communicable conditions, with 72-hour or immediate deadlines

24,624

Cancer notifications in 2024 — one registry among several

~16,400

Hours of manual abstraction a year for cancer registration alone. Derived — see below

Estimate — our calculation, not a published figure

The ~16,400 hours figure is 24,624 published cases multiplied by an assumed 40 minutes of registrar time per case. The case count is published; the per-case time is unvalidated, and measuring it is a priority at our pilot site.

The product

What Dataveoli does

An extraction sheet goes in. A completed, verifiable dataset comes out.

We install an AI abstraction engine inside the hospital’s own environment, connected to its Hospital Information System. The workflow mirrors what the staff who carry the reporting duty already do.

  1. Define the dataset

    Upload an extraction sheet — a registry data dictionary, a quality indicator specification, or a research protocol. Registry templates ship pre-built.

  2. Define the cohort

    Select patients by date range, diagnosis, service, or any criteria available in the hospital system.

  3. The system reads the charts

    Each record is processed inside the hospital’s environment, and every variable in the sheet is populated.

  4. Review what needs review

    Confident values are presented as final. The rest are flagged; the reviewer opens the cell, sees the passage, and confirms or corrects.

  5. Export and submit

    The dataset leaves in the format the destination registry requires, with a complete audit record attached.

A walkthrough of the working prototype, three and a half minutes. Every record shown is synthetic.

Why the name

The record already contains the answer.

Reading it is the work.

Every value keeps the passage it came from.

Where it is unsure, it says so.

The output

Why a hospital can trust the output

The hard problem is not extraction. It is giving a registrar a defensible reason to sign off.

A registrar signs a submission. A quality director signs an indicator return. An IRB accepts or rejects a dataset. None of them can accept a figure whose origin they cannot inspect.

Cell-level source linking

Every extracted value carries a pointer to the specific passage in the specific document it came from. A reviewer verifies a value in seconds instead of re-reading a chart.

Admission pH

Emergency department note · 14 March · 03:12

…patient obtunded on arrival, Kussmaul respiration noted. Arterial blood gas showed pH 7.08 with bicarbonate 6 mmol/L and an anion gap of 29. Insulin infusion commenced…

Select the value to reveal its source. Illustration, synthetic record.

Calibrated flagging

Where confidence in a value falls below threshold, the cell is flagged rather than filled silently. Human attention is spent where it changes the answer.

Immutable audit trail

Every extraction run, confirmation and correction is logged with user and timestamp. This makes abstraction more transparent than the manual process it replaces, in which a registrar’s judgement leaves no record at all.

The design consequence

Because we deploy models that run inside a hospital rather than the largest frontier models, we scope the product to what such models do reliably: extracting precisely defined variables from clinical text, never open-ended clinical inference. The flagging layer is what makes that scoping safe, and is therefore core to the product rather than a convenience.

Data residency

Why this has to be built here

Data residency is the entry condition for this market, and the reason the category is open.

The established players in AI chart abstraction are US-hosted. Saudi regulation governs transfer of personal data outside the Kingdom under the Personal Data Protection Law and its SDAIA transfer regulation, and health data is among the most tightly held categories. A system processing Saudi patient records must run where those records already are — an architectural requirement, not a marketing position.

System architecture. Patient-level data is processed entirely within the hospital’s boundary. Only the submission the hospital is already obliged to make leaves it, and flagged values are routed to a human reviewer inside that boundary.
  • On-premise. Hospital-owned hardware inside the hospital network. Highest assurance; higher capital cost and longer procurement.
  • In-Kingdom cloud. Sovereign cloud infrastructure located in Saudi Arabia. The same regulatory position on residency, faster to deploy, no capital outlay.

Both options run the same software from a single codebase.

Stage

Where we are

Stated plainly, because the stage matters more than the pitch.

  • A working prototype. It runs end to end: create an extraction sheet, run it, inspect any cell to reveal its source in the patient file, confirm or correct flagged values.
  • Clinical supervision in place. The clinical validation of this work is supervised at the institution where its accuracy will be measured.
  • Pre-incorporation and self-funded. No external funding has been raised. The prototype was built without external capital.
  • Looking for one pilot institution. A hospital willing to host a supervised deployment and let us measure our accuracy against its own abstraction.

What we are candid about

We are a young team without an operating company behind us, and our accuracy claim currently rests on published evidence for the model class, not yet on our own validation study. That study — measured against human abstractors on Saudi clinical documentation — is the next thing we build, and we have deliberately placed it before the commercial push rather than after it.

The people

Team

Who is building this, and who is supervising the clinical work.

Founding team

Chief Executive

Ali Alhakeem

Product & clinical workflow

Abdulaziz Alessa

Clinical validation

Ali Alkhars

Regulatory & partnerships

Ali Al Alwan

Clinical supervision

The clinical validation of this work is supervised by a clinician at the institution where its accuracy will be measured.

Dr. Sajjad AlHaddad

Diabetes Fellow, King Fahad Medical City · Senior Specialist, Family Medicine

Contact

Request demo access

The demonstration uses entirely synthetic records. Tell us who you are and we will send access.

Please enter your name.

Please enter a valid email address.

Please enter your institution.

Please enter your role.

Access links are sent by hand, usually within a day.

That did not send. The form service did not accept the request. Please email us directly — the message below is prefilled.

[email protected]

Request received

Access links are sent by hand, usually within a day. If it is urgent, email [email protected].