(CASE STUDY · 07)

MINDSCOPE

  • Data Cleaning
  • Machine Learning
  • Streamlit
  • UX | UI
  • Docker

MindScope is a mental wellbeing self-check built on 2 standard questionnaires, PHQ-9 for depression and GAD-7 for anxiety. A scikit-learn model reads both scores with 10 more answers, including age, sleep, stress and work, and returns a Low, Moderate or High risk band with advice to match.

ROLE
Data, ML & UI
TIMELINE
5 Months
YEAR
2026
TEAM
Solo

(MY ROLE)

  • Cleaned and profiled a synthetic mental health dataset in a notebook.
  • Trained 3 classifiers on 12 features and shipped the best one, a logistic regression, with its encoders.
  • Designed and built the app twice: one long form in February, a 4-step wizard in June.
  • Wrote the Dockerfile, with a build step that puts the share-card tags where crawlers can read them.
(WHY IT EXISTS)
PHQ-9 and GAD-7 are short and standard, so a self-check built on them takes minutes. The goal was to take the analysis out of the notebook and into an app.
(THE FIRST VERSION DID TOO MUCH AT ONCE)
Version 1 put all 16 questions and the lifestyle fields on one long page. On the results page, the advice sat folded inside expanders.
(WHAT CHANGED)
In June the form became a 4-step wizard with softer answer labels and a sage and clay palette. Version 1 stopped with an encoding error when gender was Other or Prefer not to say. Options the encoders never saw now fall back to a known class.

Process

  1. 01

    Data & Scales

    Cleaned the Kaggle Global Mental Health Dataset 2025 in a notebook: 2,500 records, 15 columns, 10 countries. Scores map to the standard PHQ-9 and GAD-7 severity bands.

  2. 02

    Label & Model

    The risk label is the mean of the two scores: under 5 Low, under 12 Moderate, otherwise High. Encoded 7 categorical columns and trained 3 classifiers on a stratified 80/20 split, keeping the best test score.

  3. 03

    First Release

    Wrapped the saved model and encoders in a Streamlit app that loads them once per process. The app and a Dockerfile were committed on the same day in February.

  4. 04

    Redesign

    Rebuilt the interface in June as a 4-step wizard that keeps each answer across steps. Phone layouts came the same day.

How It Works

Answer 16 questions over 4 steps. The app adds up the PHQ-9 and GAD-7 scores, encodes the other answers and asks a saved model for a risk band. Every answer stays in the session. Nothing is written down.

The Check-In

The subject is sensitive, so the goal was reflection, not alarm. Warm sand, sage and clay, serif headings, and answers in plain words: Not really, Here & there, Quite a lot, Nearly always. The chips shade from sage to clay as an answer grows.

In The Detail

A running score sits under each questionnaire and moves with every answer. The result leads with the band and the model's confidence, then the scores. Below, the same question in version 1 and version 2.

(SYSTEM 01 · RISK MODEL)

Risk Model

A logistic regression, kept from 3 classifiers and cached with st.cache_resource. It reads 12 numbers and returns Low, Moderate or High, with a probability for each.

(NOT ONE RECORD IS A REAL PERSON)
The Kaggle dataset is generated, not collected from patients. The model has never seen a real person's answers.
(THE LABEL IS A FORMULA OF 2 INPUTS)
Both scores that define the risk label are also among the 12 features. So the model mostly relearns the formula, and its test score reflects that.

(SCOPE)

questions across 2 scales
16
features in every prediction
12
records in the synthetic dataset
2,500

(SYSTEM 02 · GUIDANCE)

Guidance

Every band opens its own advice: 50 tips in 14 categories, 12 for Low, 17 for Moderate and 21 for High. Each band has its own colour and banner. A High result opens with Immediate Actions, then treatment, crisis lines, daily wellness and self-care; a Low result keeps to activity, sleep, mental wellness and lifestyle. Every result ends with a reminder that this is not a diagnosis.

  1. The three bands

  2. A High result

  3. Treatment and self-care

  4. A Low result

Shipped

The June version is in the repo. The Render service is suspended, so the source is the way to run it.