SMLEREVISE Technical Architecture
A transparent look at the systems behind SMLEREVISE — the IRT-lite engine that reproduces the SCFHS 200–800 scaled score, a Medical Knowledge Graph built on Saudi MOH protocols, a Google Embedding 2 semantic deduplication pipeline, and an adaptive SM-2 spaced-repetition system. Everything on this page is part of our public verification record.
Platform at a Glance
The core specification of the SMLEREVISE platform in one table. Every row is expanded in the sections below.
| Active question corpus | 5,000+ calibrated MCQs & clinical vignettes |
|---|---|
| Concept coverage | ~4,500 unique concept-clinical patterns across the SCFHS Blueprint |
| Scoring engine | IRT-lite on the SCFHS 200–800 scale |
| Passing threshold | 560 (scaled score) |
| Deduplication | Google Embedding 2 — 768-dim vectors, 0.85 cosine threshold |
| Spaced repetition | Adaptive SM-2 with weak-category injection |
| Content alignment | Saudi MOH protocols & SCFHS Blueprint |
| Update cadence | Within 7 days of SCFHS blueprint modifications |
| Medical review | Dual review by licensed Saudi physicians |
| Security | AES-256 at rest, TLS 1.3 in transit |
Medical Knowledge Graph (MKG)
The Medical Knowledge Graph is not a diagram — it is the relational data structure every SMLEREVISE product is built on. Each node (disease, investigation, drug, procedure) is linked to its neighbours through the causal and therapeutic relationships exactly as they appear in Saudi Ministry of Health protocols, so revising one concept pulls in the Saudi-specific context the SMLE expects you to know.
Take a single symptom. The cough node alone connects to 47 clinical entities in the graph — and one of those paths carries the exact local context the SCFHS likes to test: from presentation, through Saudi epidemiology, to the Ministry of Health's first-line treatment protocol.
The same path in text:
- Cough (symptom) — a clinical presentation linked to 47 nodes in the knowledge graph
- Tuberculosis (disease) — flagged by the cough node, with Saudi prevalence data attached: 15.2 per 100,000 population (WHO 2024)
- MOH first-line protocol — INH + RIF + PZA + EMB for 6 months, the first-line regimen in Saudi Arabia
Semantic Deduplication Engine (Google Embedding 2)
SMLEREVISE is the only SMLE platform that uses Google's Embedding 2 model to compute semantic similarity between questions. Every new question is embedded as a 768-dimensional vector and compared against the entire corpus; if the cosine similarity against any existing question exceeds 0.85, the item is flagged as a duplicate and routed to editorial review for merging or removal.
How a question gets screened
- Embedthe question is converted into a 768-dimensional vector with Google Embedding 2
- Comparecosine similarity is computed against every existing question
- Flaganything above 0.85 is marked as a likely duplicate
- Revieweditors merge or remove it; the same scenario with swapped lab values is caught automatically
Why it matters: redundancy is the hidden tax of big question banks — answering the same concept ten times teaches you nothing new the ninth time. Deduplication keeps ~4,500 unique concept-clinical patterns covering the entire SCFHS Blueprint, with zero filler. For reference: UWorld, the gold standard for the hardest medical exams, runs on ~4,000 questions.
IRT-lite Scaled Scoring Engine (200–800)
SMLEREVISE is the only platform that simulates the SCFHS scaled scoring system. Instead of a raw percentage, every Grand Mock is scored with an Item Response Theory (IRT) engine at Lite level: each question carries calibrated difficulty, discrimination, and guessing parameters, and your responses are converted into an ability estimate θ — the same psychometric logic behind the real 200–800 score report.
P(θ) = c + (1 − c) / (1 + e−a(θ − b))The ability estimate is then transformed to the SCFHS scale — mean 500, standard deviation 100, bounded between 200 and 800 — with the passing standard at 560. A SMLEREVISE score therefore reads exactly like the score report you will receive on exam day, not an inflated percentage.
Adaptive SM-2 Spaced Repetition
Flashcards in SMLEREVISE run on an optimized SM-2 algorithm. Each review reschedules the card by multiplying the previous interval by an ease factor between 1.3 and 2.5, which the algorithm raises or lowers based on how you actually perform — not on a fixed calendar.
I(n) = I(n−1) × EFTwo SMLE-specific adaptations sit on top of classic SM-2. First, intelligent injection: when your accuracy in a category drops below 70%, the system automatically injects 5–10 high-yield cards targeting that weakness. Second, exam-date calibration: intervals are tuned against the SMLE date you set, so high-yield material resurfaces before you sit the exam.
Medical Review Protocol
Every piece of medical content on SMLEREVISE passes through a documented review protocol before it reaches you, and stays under review after publication. The protocol exists to guarantee one thing: compliance with current Saudi Ministry of Health protocols and the SCFHS Blueprint.
- First reviewa licensed Saudi physician checks the content for clinical accuracy
- Protocol verificationthe item is compared against current Saudi MOH protocols
- Drug verificationevery drug is confirmed as available in Saudi Arabia and holding first-line status where claimed
- Second reviewhigh-risk content (dosages, interactions, contraindications) is reviewed independently again
- Documentationeach review is logged with reviewer name, date, and content version
Security & Data Protection
Candidate data and medical content are protected with the same rigor as the question bank itself. SMLEREVISE encrypts data end to end, verifies every session, and keeps exam content safe from capture during mock exams.
Technical FAQs
It uses Google's Embedding 2 model to convert every question into a 768-dimensional vector. If the cosine similarity exceeds 0.85 with any existing question, it is flagged as a duplicate and reviewed for merging or removal.
Ready to Prepare the Right Way?
SMLEREVISE is the only platform built from the ground up for the SMLE — with real SCFHS-scaled scoring, Saudi MOH-aligned content, and Sina for instant concept clearing without leaving the platform.