Belva.
Back to Works
Data EngineeringAI Generated Dummy Data

Youth Mental Health Early Warning System (DKI Jakarta)

End-to-End Distributed Big Data Pipeline, Medallion Lakehouse Architecture & GenAI Analytics on Databricks

#Databricks#PySpark#Delta Lake#Unity Catalog#Medallion Architecture#Lakeflow Workflows#GenAI Genie Agent#Healthcare Analytics#SQL Dashboards

01. Problem Identification

Amid rapid digitalization in DKI Jakarta, the younger generation faces escalating anxiety and depression due to excessive screen exposure, late-night social media scrolling, and high academic pressure. Traditionally, public health responses have been reactive—the government only becomes aware of a crisis when patients are already in severe condition. The lack of integration between self-reported telemetry logs, clinic clinical assessments, and counseling history leads to delayed early interventions.

PAIN POINT 01

Multi-Silo Data Fragmentation: 300,000 records of device telemetry, self-reported psychometric questionnaires, and counseling logs are scattered across different formats without a single source of truth.

PAIN POINT 02

Sensor Anomalies & Medical Privacy Standards: High sensor noise (screen time > 20 hours/day) and the obligation to protect medical privacy records (HIPAA/GDPR) to prevent youth identity leaks.

PAIN POINT 03

Lack of Spatial Predictive Systems: Absence of high-risk regional mapping and forecasting systems to anticipate seasonal surges (e.g., during school exam periods).

02. The Solution

Built a unified Early Warning System using the Databricks Lakehouse ecosystem. The Medallion Architecture cleans and integrates 300,000 data records in a distributed manner using PySpark, secures privacy via SHA-256 cryptography, automates data flows with Lakeflow Jobs, and provides a Databricks Genie AI agent for natural language exploration along with interactive BI dashboards featuring 30-day predictive projections.

Medallion Architecture & Quarantine Pattern

Controlled ingestion in Bronze with StructType, automated isolation of anomalous records into a Quarantine table, and smart logic cleansing in Silver.

Cryptographic PII Masking (HIPAA Compliance)

Transformation of user_id into consistent one-way SHA-256 hashes to maintain inter-table relationships without exposing patient personal identities.

Automated Lakeflow Workflows & Validation

Execution of multi-stage parameterized SQL batches using dbutils.widgets with automated integrity validation tests (PASS/FAIL).

Databricks Genie AI Agent & Semantic Layer

Provisioning a semantic Business View, medical terminology dictionary, domain instructions, and 8 trusted benchmark queries for instant hallucination-free analytics chat.

Predictive 30-Day Forecasting & BI Dashboard

Interactive visualization with 5 KPI cards, regional morbidity maps, screen time vs sleep loss correlations, and projected healthcare facility needs via AI_FORECAST.

03. Crafting & Engineering

System Architecture Flow

#01Raw Multimodal Sources3x 100K CSVs (Telemetry, GAD-7/PHQ-9 Assessments, Counseling)
Unity Catalog Volumes
#02Bronze Lakehouse LayerRaw Delta ingestion with StructType schema enforcement & audit trail
Delta Lake / ACID
#03Silver Layer & DQ EngineOutlier quarantine, PII SHA-256 masking, timestamp parsing & imputation
PySpark Distributed
#04Gold Data MartsMulti-table window joins, KPI aggregates, cross-table assertions
Databricks SQL Warehouse
#05Lakeflow OrchestrationParameterized batches, dependency graphs & automated repair runs
Databricks Workflows
#06Genie AI & Live BI DashboardSemantic Business View, NLP Text-to-SQL & 30-day AI Forecasting
Databricks Genie / BI

Medallion Lakehouse & Automation Pipeline

Bronze Layer

Raw Data Onboarding & Schema Enforcement

Unity Catalog Volumes / Delta Format

Storage of raw CSVs in Unity Catalog Volume (raw_dataset), strict StructType schema enforcement, recording provenance metadata (_metadata.file_path, ingested_at), and compilation into ACID-guaranteed Delta Lake.

Silver Layer

Data Quality Engine, Quarantine & Masking

PySpark 3.4 (DataFrame API)

Data profiling, separating screen time anomalies (>18 hours) to Quarantine Delta Table, capping late-night social media, smart imputation of null values for sleep index, and PII identity masking with SHA-256 Hashing.

Gold Layer

Curated Business Data Marts & Aggregations

Databricks SQL / Delta Lake

Multi-table relational joins based on user_id_masked with ROW_NUMBER() window functions, creation of regional & stressor analytical data marts, and data reconciliation assertions to guarantee data integrity.

Orchestration

Lakeflow Automation & Validation Workflows

Databricks Lakeflow Jobs

Pipeline automation using Parameterized Notebooks (dbutils.widgets), 4-stage SQL Batch processing, dependency DAGs, validation assertions (PASS/FAIL), and automated Repair Run mechanisms.

GenAI & BI

Databricks Genie AI Agent & EWS Dashboard

Databricks Genie + Databricks SQL

Provision of Semantic Business View for GenAI Agent, clinical synonym dictionary, Trusted Queries, benchmark testing suite, and EWS visualization complete with 30-day AI_FORECAST.

Multimodal Datasets (300,000 Total Rows)AI-Generated Dummy Data

telemetry_digital_youth.csv100,000 rows

Daily activity telemetry logs: screen time, late-night social media time (22:00–04:00), sleep quality index, exercise duration, and device OS type.

CSV → Bronze Delta Table
assessment_psychological_youth.csv100,000 rows

Standardized medical questionnaire results: GAD-7 anxiety scores (0–21), PHQ-9 depression scores (0–27), major stress triggers, and age range categories.

CSV → Bronze Delta Table
counseling_intervention_peer.csv100,000 rows

Peer tele-counseling intervention logs: satisfaction scores, message sentiment analysis scores, session modalities, and clinic/hospital referral recommendation status.

CSV → Bronze Delta Table

Technology Selection

Databricks Lakehouse & Delta LakeProvides distributed computing scalability, ACID transactions, schema enforcement, and Time Travel features for medical data auditing.
PySpark 3.4 (DataFrame API)Processes parallel transformations for 300,000 rows of telemetry data, window function queries, and in-memory feature engineering.
Unity Catalog (Volumes & Governance)Manages centralized access control, non-tabular data governance in raw_dataset, and data lineage from upstream to downstream.
Databricks Genie AI Agent (AI/BI)Democratizes clinical insights access for non-technical policymakers through Natural Language Text-to-SQL with a Semantic Layer.
Lakeflow Workflow / Databricks JobsAutomates multi-task dependency scheduling, dynamic parameter passing, and automated partial repairs (Repair Run).
Databricks SQL & AI_FORECASTBuilds interactive KPI visualizations for the Early Warning System and projects estimated crisis intervention needs for the next 30 days.

Key Engineering Challenges Solved

  • Handling timestamp format heterogeneity and sensor logic anomalies (late-night social media use > total screen time) without breaking daily time-series data integrity.
  • Implementing a data quarantine architecture to isolate anomalous data (screen time > 18 hours) to prevent bias in calculating regional average morbidity.
  • Maintaining patient medical record privacy according to HIPAA/GDPR standards via one-way masking (Cryptographic Hashing SHA-256) on user_id while preserving inter-table relationships.
  • Eliminating AI hallucination risks in the Databricks Genie Agent by designing a semantic Business View, clinical synonym dictionary, and a collection of Trusted Queries.
  • Performing Data Reconciliation using automated cross-layer assertion checks to guarantee zero data loss after multi-table joins.

04. Interactive Lakehouse Dashboard

Live embedded Databricks SQL Early Warning System with interactive filters and predictive analytics.

Open in Databricks

Initializing Databricks AI/BI Session...

Menghubungkan ke Databricks Workspace & rendering visualisasi

5 Key Indicator Cards

Memantau 23.562 remaja aktif, 5.853 kasus rujukan (24.84%), rata-rata durasi layar 8.01 jam, dan rata-rata skor depresi PHQ-9 (13.8) di DKI Jakarta.

Disparitas Wilayah & Stresor

Jakarta Selatan dan Timur mencatatkan tingkat rujukan tertinggi. Pemicu stres utama didominasi beban akademik (31.8%) dan konflik keluarga (18.6%).

Proyeksi 30 Hari (AI_FORECAST)

Fitur prediktif built-in Databricks memproyeksikan volume kedaruratan krisis untuk perencanaan kapasitas intervensi faskes & konselor sebaya secara antisipatif.

05. Databricks Genie AI Agent

Interaksi kecerdasan buatan berbasis Natural Language Text-to-SQL dengan 2 mode eksplorasi.

Open in Databricks

Databricks Genie AI

Connected

Semantic Text-to-SQL Lakehouse Intelligence

Perhatian: Jawaban pada mode Public Live Chat ini menggunakan pemanggilan Text-to-SQL API langsung (non-agentic), sehingga tingkat akurasinya tidak seoptimal versi Full Native Space yang memiliki penalaran analitis bertahap.

Halo! Saya adalah Databricks Genie AI Agent (Public Demo) untuk Youth Mental Health Early Warning System. Anda dapat menanyakan data statistik atau tren klinis dengan bahasa sehari-hari.

⚠️ Catatan: Mode Public Live Chat ini terhubung via Text-to-SQL API langsung (non-agentic), sehingga tingkat akurasi dan konteksnya tidak seakurat versi Full Native Space yang dilengkapi penalaran agentic bertahap.

09:19 AM
Saran Pertanyaan: