Technical Documentation

SMLEREVISE Technical Architecture

A transparent look at the systems behind SMLEREVISE — the IRT-lite engine that reproduces the SCFHS 200–800 scaled score, a Medical Knowledge Graph built on Saudi MOH protocols, a Google Embedding 2 semantic deduplication pipeline, and an adaptive SM-2 spaced-repetition system. Everything on this page is part of our public verification record.

  • Medically reviewed by Dr. M. Salar Raza — last reviewed 14 May 2026
  • Updated within 7 days of SCFHS blueprint changes

Platform at a Glance

The core specification of the SMLEREVISE platform in one table. Every row is expanded in the sections below.

Active question corpus5,000+ calibrated MCQs & clinical vignettes
Concept coverage~4,500 unique concept-clinical patterns across the SCFHS Blueprint
Scoring engineIRT-lite on the SCFHS 200–800 scale
Passing threshold560 (scaled score)
DeduplicationGoogle Embedding 2 — 768-dim vectors, 0.85 cosine threshold
Spaced repetitionAdaptive SM-2 with weak-category injection
Content alignmentSaudi MOH protocols & SCFHS Blueprint
Update cadenceWithin 7 days of SCFHS blueprint modifications
Medical reviewDual review by licensed Saudi physicians
SecurityAES-256 at rest, TLS 1.3 in transit

Medical Knowledge Graph (MKG)

The Medical Knowledge Graph is not a diagram — it is the relational data structure every SMLEREVISE product is built on. Each node (disease, investigation, drug, procedure) is linked to its neighbours through the causal and therapeutic relationships exactly as they appear in Saudi Ministry of Health protocols, so revising one concept pulls in the Saudi-specific context the SMLE expects you to know.

Take a single symptom. The cough node alone connects to 47 clinical entities in the graph — and one of those paths carries the exact local context the SCFHS likes to test: from presentation, through Saudi epidemiology, to the Ministry of Health's first-line treatment protocol.

The cough pathway in the Medical Knowledge Graphindicatestreated bylocal epidemiologyCoughSymptomTuberculosisDiseaseMOH First-Line ProtocolINH + RIF + PZA + EMB6 months — first-line in KSASaudi TB Prevalence15.2 per 100,000 · WHO 2024
Figure 1 — The cough pathway as stored in the MKG: presentation → local epidemiology → first-line treatment.

The same path in text:

  1. Cough (symptom) — a clinical presentation linked to 47 nodes in the knowledge graph
  2. Tuberculosis (disease) — flagged by the cough node, with Saudi prevalence data attached: 15.2 per 100,000 population (WHO 2024)
  3. MOH first-line protocol — INH + RIF + PZA + EMB for 6 months, the first-line regimen in Saudi Arabia

Semantic Deduplication Engine (Google Embedding 2)

SMLEREVISE is the only SMLE platform that uses Google's Embedding 2 model to compute semantic similarity between questions. Every new question is embedded as a 768-dimensional vector and compared against the entire corpus; if the cosine similarity against any existing question exceeds 0.85, the item is flagged as a duplicate and routed to editorial review for merging or removal.

How a question gets screened

  1. Embedthe question is converted into a 768-dimensional vector with Google Embedding 2
  2. Comparecosine similarity is computed against every existing question
  3. Flaganything above 0.85 is marked as a likely duplicate
  4. Revieweditors merge or remove it; the same scenario with swapped lab values is caught automatically

Why it matters: redundancy is the hidden tax of big question banks — answering the same concept ten times teaches you nothing new the ninth time. Deduplication keeps ~4,500 unique concept-clinical patterns covering the entire SCFHS Blueprint, with zero filler. For reference: UWorld, the gold standard for the hardest medical exams, runs on ~4,000 questions.

IRT-lite Scaled Scoring Engine (200–800)

SMLEREVISE is the only platform that simulates the SCFHS scaled scoring system. Instead of a raw percentage, every Grand Mock is scored with an Item Response Theory (IRT) engine at Lite level: each question carries calibrated difficulty, discrimination, and guessing parameters, and your responses are converted into an ability estimate θ — the same psychometric logic behind the real 200–800 score report.

P(θ) = c + (1 − c) / (1 + e−a(θ − b))
θ candidate abilitya discriminationb difficulty (−3.0 to +3.0)c guessing probability (0.25)

The ability estimate is then transformed to the SCFHS scale — mean 500, standard deviation 100, bounded between 200 and 800 — with the passing standard at 560. A SMLEREVISE score therefore reads exactly like the score report you will receive on exam day, not an inflated percentage.

Figure 2 — The SCFHS 200–800 scaled score. The passing standard (560) sits 0.6 standard deviations above the mean.

Adaptive SM-2 Spaced Repetition

Flashcards in SMLEREVISE run on an optimized SM-2 algorithm. Each review reschedules the card by multiplying the previous interval by an ease factor between 1.3 and 2.5, which the algorithm raises or lowers based on how you actually perform — not on a fixed calendar.

I(n) = I(n−1) × EF

Two SMLE-specific adaptations sit on top of classic SM-2. First, intelligent injection: when your accuracy in a category drops below 70%, the system automatically injects 5–10 high-yield cards targeting that weakness. Second, exam-date calibration: intervals are tuned against the SMLE date you set, so high-yield material resurfaces before you sit the exam.

Medical Review Protocol

Every piece of medical content on SMLEREVISE passes through a documented review protocol before it reaches you, and stays under review after publication. The protocol exists to guarantee one thing: compliance with current Saudi Ministry of Health protocols and the SCFHS Blueprint.

  1. First reviewa licensed Saudi physician checks the content for clinical accuracy
  2. Protocol verificationthe item is compared against current Saudi MOH protocols
  3. Drug verificationevery drug is confirmed as available in Saudi Arabia and holding first-line status where claimed
  4. Second reviewhigh-risk content (dosages, interactions, contraindications) is reviewed independently again
  5. Documentationeach review is logged with reviewer name, date, and content version

Security & Data Protection

Candidate data and medical content are protected with the same rigor as the question bank itself. SMLEREVISE encrypts data end to end, verifies every session, and keeps exam content safe from capture during mock exams.

Technical FAQs

It uses Google's Embedding 2 model to convert every question into a 768-dimensional vector. If the cosine similarity exceeds 0.85 with any existing question, it is flagged as a duplicate and reviewed for merging or removal.

Ready to Prepare the Right Way?

SMLEREVISE is the only platform built from the ground up for the SMLE — with real SCFHS-scaled scoring, Saudi MOH-aligned content, and Sina for instant concept clearing without leaving the platform.