Automated Phenotype Discovery
Unsupervised clustering to identify response-homogeneous subgroups

Data-Driven Stratification

Applies k-means clustering on 15 baseline variables (baseline EF, reaction time, anxiety, social preference, etc.). Identifies 4-5 distinct learner phenotypes.

Rapid Learners (Type A)

High baseline working memory, fast response time, low anxiety

✓ Benefit from accelerated progression

Careful Processors (Type B)

Slower response, high accuracy, high perfectionism

✓ Need confidence-building, avoid speed pressure

Social Learners (Type C)

High peer sensitivity, improve with competition

✓ Multiplayer/social features boost engagement

Stability-Seekers (Type D)

Prefer predictable schedules, sensitive to changes

✓ Consistent routine improves compliance

ML Pipeline

  • • PCA dimensionality reduction (15 variables → 5 components, 82% variance explained)
  • • Elbow method determines optimal k=4 clusters
  • • Silhouette score validates cohesion (avg=0.68, good)
  • • Cluster characteristics interpreted via domain expertise

Intervention Personalization

Type A gets adaptive difficulty acceleration; Type B gets accuracy feedback; Type C gets tournament modes; Type D gets fixed weekly schedule. Average ROI increase: 31%.

base44
Edit with Base44