Data Science Part-Time capstone projects batch #11

by Christina Sieber

students

We’re thrilled to celebrate the achievements of our latest graduates from Part-Time Data Science Batch #11, who have just wrapped up their Data Science journey with 4 remarkable, real-world projects. This round of final presentations showcased how data science and AI can drive tangible impact across industries, from transforming business development workflows to reinventing the way of the new market discovery. 

Take a look at how our graduates are using data science to generate insights, push boundaries, and create real-world impact. 

 

Predicting Stablecoin Depegging Risk: An Actionable Early Warning Framework

Project by: Yuqing Sun

Stablecoins have quietly grown into the $300B+ backbone of cryptocurrency transactions, logging over $33T in volume in 2025 alone. Major issuers like Tether and Circle hold massive portfolios of U.S. Treasury bills, binding token stability directly to sovereign debt markets and financial regulation. Yet when a peg breaks—whether it is Terra’s $40B collapse in 2022 or USDC slipping to $0.91 during the 2023 Silicon Valley Bank run—the resulting panic drains liquidity across both decentralised and traditional markets.

To address this systemic risk, Yuqing developed an end-to-end monitoring framework capable of warning against depegging events hours in advance and executing capital-protective exits.

 

Analytical Framework & Data Architecture

The project analyzed 9 years of 15-minute resolution panel data (2017–2026) encompassing 11 stablecoins and 3.5M+ observations. The model ingested 53 engineered features spanning six distinct data sources: exchange pricing, decentralised exchange (DEX) slippage, DeFi APYs, futures premiums, total token supply, and TradFi macroeconomic indicators.

  1. Depeg Detection: A Random Forest classifier trained on 3-fold Walk-Forward Cross-Validation (strictly preventing forward-looking bias) identified the onset of binary depegging events.
  2. Survival Time Estimation: An XGBoost Accelerated Failure Time (AFT) model predicted safe hours remaining before a peg failure. Simultaneously, a Network-LSTM deep learning architecture with attention mechanisms captured cross-token contagion dynamics, achieving a 0.965 concordance index (C-index).
  3. Regime Context: An unsupervised 4-state Hidden Markov Model (HMM) categorized market conditions into four distinct stress regimes: Calm, Active, Stress, and Depeg.


Key Findings & Trading Execution

  • Lead Time by Token Architecture: Algorithmic stablecoins (e.g., UST, USDD) yielded 100% event recall (62/62 events caught) with a 34–44 hour early warning window. Over-collateralized assets (e.g., DAI) achieved 90–100% recall with 28–45 hours of lead time. Fiat-backed tokens showed lower recall (29–42%) because their primary failure triggers stem from off-chain banking events invisible to on-chain market data.
  • Macro Triggers vs. Liquidity Fuel: TradFi macro signals (VIX, financial sector ETF volatility) proved to be the top individual feature drivers, but crypto microstructure features dominated cumulative feature importance (51.7% vs. TradFi's 29.4%). Combining HMM stress regimes with SHAP values revealed that equivalent DEX selling pressure is 3 to 5 times more likely to trigger a depeg during a high-stress market regime than during calm conditions.
  • 96% Reduction in Execution Friction: Standard binary alert models triggered 389 trades during the 2024–2026 backtest period due to high-frequency market noise. By replacing binary thresholds with a Restricted Mean Survival Time (RMST) continuous exit strategy (triggering position reductions when predicted safe hours dropped below 18), execution overhead dropped to 13 trades while maintaining ~41% depeg downside protection.
 

Unlocking Field Intelligence: Automated Extraction of Crop Damage Notes

Project by: Lorin Semela

When severe weather strikes Swiss vineyards, loss adjusters from agricultural insurer SHGroup inspect affected fields and log their findings into custom tablet software. While adjusters record standardized data, critical nuances—such as exact counts of damaged flower buds versus total evaluated buds, or the presence of an undamaged frost reserve branch—are recorded inside a single free-text comment box.

Because these notes were written in unstructured German and French, valuable information remained inaccessible to underwriting and R&D teams. Lorin designed an automated natural language processing pipeline to extract structured metrics directly from raw expert commentary without altering the adjusters' field workflow.


Pipeline Architecture & Privacy Controls

Because crop assessment data contains sensitive business information, standard external LLM APIs were ruled out. Lorin built an end-to-end pipeline hosted entirely within Microsoft Fabric using a locally deployed open-source language model.

  • Data Cleaning: Regular expressions stripped repetitive boilerplate, signature blocks, and noise from incoming comments.
  • Benchmark Dataset: Lorin manually annotated a ground-truth dataset of 150 real-world expert notes to evaluate extraction performance across multiple candidate models.
  • Model Selection: Qwen2.5-1.5B-Instruct was selected for its compact local resource usage, strong multilingual comprehension (German and French), and reliable structured output formatting.
  • Task Splitting: Initial experiments using a single prompt produced inconsistent numerical extraction. Splitting the prompt into focused sub-tasks—separating categorical checks from exact numeric counting—substantially improved accuracy.

 

Field Extraction Results

 

Task

Performance Metric

Frost damage detected

98.7% Overall Accuracy (94.1% F1)

Reserve branch present

98.7% Overall Accuracy (93.3% F1)

Count of damaged buds

93.8% Exact Match Rate

Total buds evaluated

95.7% Exact Match Rate

By converting qualitative field commentary into clean tabular records, SHGroup gained immediate, privacy-compliant access to historical and ongoing assessment data, improving risk modeling without adding administrative work for field experts.
 

Hospitality Revenue Protection: Profit-Driven Cancellation Forecasting

Project by: Richard Korcz

In the hotel industry, inventory is entirely perishable: an unbooked room on any given night is revenue lost forever. With the average canceled reservation costing hotel operators €296 in lost booking value, relying on static cancellation policies leads to substantial margin leakage.

Richard engineered an end-to-end machine learning solution designed to identify high-risk reservations months ahead of check-in, enabling operators to execute targeted customer-retention campaigns before the booking is lost.

 

The Technical Edge

The project analyzed over 33,000 guest reservations, scrubbing incomplete records and using Box-Plot filtering to remove revenue extreme outliers.

  1. Behavioral Feature Engineering: Engineered key ratio metrics including Revenue per Guest, total length of stay, and cumulative financial exposure.
  2. Class Imbalance Resolution: Because non-canceled bookings represent 68% of raw data, the training set was balanced using random undersampling to prevent decision bias toward the majority class.
  3. Key Risk Profiling: Baseline logistic regression and automated PyCaret multi-model pipelines identified early lead time (+1.48 coefficient), high room rates (+0.33), and online booking friction (+0.39) as the primary indicators of cancellation risk. Conversely, special guest requests (-1.19) and parking reservations (-0.29) signaled highly committed bookings.
​​The Financial Asymmetry

Standard machine learning models often optimize purely for overall accuracy. Richard’s model was evaluated explicitly against asymmetric business economics:

  • Cost of a Missed Cancellation (False Negative): In the evaluation sample, 696 cancellations were missed by the model. At €296 per lost room, unmitigated cancellations cost €206,016 in lost revenue.
  • Cost of a Targeted Incentive (False Positive): Issuing a €30 promotional incentive (~10% discount) to 1,511 flagged bookings costs €45,330.

Optimizing model hyper-parameters specifically for Recall ensures that high-risk bookings are detected early. Because issuing a promotional discount costs less than one-fifth of an unmitigated room cancellation, proactive retention spending delivers immediate, measurable ROI for hospitality management.
 

The Meal Counter: Automating Smart Fridge Inventory & Loss Prevention

Project by: Alejandro Soares, Christopher Lan

FELFEL has transformed office dining by placing smart vending fridges filled with fresh meals, snacks, and drinks directly into modern workplaces. However, relying on unmonitored inventory opens vulnerabilities in stock accuracy and revenue security. Theft and untracked removals directly cause financial losses, while inventory discrepancies leave customers frustrated when expected items are missing.

To secure inventory without adding friction to the customer experience, we built The Meal Counter—an automated computer vision system that tracks fridge inventory in real-time using embedded edge hardware.
 

Hardware Integration & Detection Architecture

The pipeline pairs low-cost edge hardware with high-capacity object detection models to capture stock movements as users interact with the fridge.

  • Hardware & Data Acquisition: An onboard Raspberry Pi and Raspberry Pi Camera mounted inside the smart fridge unit capture live video feeds of the shelves during customer interactions.
  • Software Stack: Product image annotation was managed in Label Studio, while object detection models were trained and deployed using the Ultralytics framework.
  • Core Model Architecture: The system leverages a YOLOv11l deep learning model for real-time bounding box detection.
  • Dual-Module Workflow:
    1. Object Detection: Identifies specific product categories and meal types as items move.
    2. Counter Module: Continuously tracks bounding box trajectories to compute net changes, logging exactly which items were removed or put back.


Model Development: Navigating Human Interaction

Detecting products in real-world vending environments introduces unique edge cases, particularly when human hands and heads enter the camera frame during selection. The team executed three experimental iterations to refine classification accuracy:

Model Iteration

Dataset Strategy

Observed Issue

Technical Solution

V0: Initial Background Model

Trained with standard background images.

Model erroneously classified customers' heads as meal products.

Requires explicit human component labeling.

V1: Head Integration Model

Dataset updated to include labeled human heads.

Model misclassified user hands as heads.

Shifted strategy to designate heads as background objects.

V2: Calibrated Background Model

Labeled heads explicitly as background regions.

Produced an increased rate of false-positive product detections.

Highlighted the need for spatial counting zones and tighter confidence thresholds.

Benchmark model evaluations achieved high accuracy metrics of up to 98% during testing. Video trial runs successfully demonstrated the system's ability to identify products and track items taken or returned in real time. Ongoing calibration focuses on fine-tuning confidence thresholds to eliminate false positives and "ghost" item detections caused by varying production lighting. 


Business Impact & Strategic Vision

By replacing manual stock checks with automated computer vision, The Meal Counter creates an immediate layer of financial protection for smart hospitality operations. Real-time tracking reduces inventory shrinkage, safeguards margins, and ensures accurate stock status for customers.

The ultimate goal is a fully generalized vision pipeline capable of automatically detecting, classifying, and settling purchases across FELFEL's entire product catalog—delivering seamless, friction-free catering to modern workplaces.

Interested in reading more about Constructor Nexademy and tech related topics? Then check out our other blog posts.

Read more
Blog