uScheme Interpreter Optimization
Optimized closure creation by analyzing free variables across the language AST and retaining only bindings a lambda could actually reference.
Computer Science + Mathematics @ Tufts University
I build and optimize systems, data pipelines, and machine-learning models — especially where performance, evaluation, and real-world constraints matter.
Selected work
Chosen for depth, engineering decisions, measurable outcomes, and breadth across systems, ML, and product work.
Optimized closure creation by analyzing free variables across the language AST and retaining only bindings a lambda could actually reference.
Profiled and reworked an unfamiliar 32-bit virtual-machine implementation, changing data representations and execution structure to remove overhead and improve locality.
Built a text-classification pipeline combining character 3–5-gram TF-IDF with 28 readability and stylometric features.
Combined clinical metadata with engineered image features, using patient-grouped evaluation and class weighting on a heavily imbalanced dataset.
Indexed text across recursive directory trees with a custom templated hash table using linear probing and rehashing, then exposed case-sensitive and case-insensitive CLI search.
Source kept private in accordance with course policy.
Designed a modular, reversible visual modernization layer for Tufts’ legacy PeopleSoft SIS while preserving existing behavior and navigation.
Case studies
These are intentionally implementation-focused without publishing restricted course solutions.
01 / Programming languages + performance
The interpreter created closures by carrying the surrounding environment. That is semantically safe, but potentially wasteful when a lambda references only a small subset of those bindings.
Implemented recursive free-variable analysis over the expression tree, including binding forms, nested lambdas, mutation, control flow, and function application. Closure creation then filtered the captured environment to bindings proven relevant to the lambda.
Added targeted checks around free-variable behavior, integrated the analysis at closure creation, and benchmarked the modified interpreter against the original source implementation.
Runtime fell from 5.14 seconds to 1.35 seconds on the assignment benchmark — a ~74% reduction. The project demonstrates interpreter internals, lexical scope, recursive AST analysis, and performance optimization without changing language semantics.
Source kept private in accordance with course policy.
02 / Systems + performance engineering
Start with an unfamiliar Universal Machine implementation and make it faster without changing its instruction semantics. The initial design relied on higher-overhead container abstractions for frequently accessed runtime state.
Progressively moved segment and reusable-ID representations from Hanson sequences to UArrays and then flatter arrays, and consolidated execution paths to reduce indirection and improve locality.
A consolidation pass unexpectedly exhausted stack space after moving state into local scope; the fix was to allocate the large runtime state on the heap. Callgrind identified instruction dispatch as the dominant hot path, and generated assembly confirmed the compiler lowered the opcode switch to a jump table with frequently used state kept largely in registers.
Performance work as an evidence loop: profile first, change representation based on a bottleneck hypothesis, debug secondary effects, then inspect machine-level behavior rather than assuming the compiler did what the source suggested.
Team project. Source kept private in accordance with course policy.
03 / ML evaluation + NLP
Classify literary passages into two reading-level groups from text plus computed readability and stylometric metadata.
The final pipeline used character-level 3–5-gram TF-IDF alongside 28 numerical features, standardized the numeric branch, and trained logistic regression.
Ordinary 10-fold validation produced a mean AUROC around 0.82, but inspection showed that passages from the same book could land in both training and validation folds. That made the model partly answer “have I seen this work before?” instead of the intended reading-level question.
Rebuilt validation around groups defined by author + title using stratified group folds, so no literary work crossed the train/validation boundary. The lower, more variable grouped scores were treated as more trustworthy than the prettier leaked result.
Course project. Source kept private; no assignment solution code is published here.
04 / Applied ML + multimodal features
Predict a binary skin-lesion target from patient metadata plus smartphone images. The training data contained 1,178 examples and a severe 92.4% positive-class imbalance.
Combined 55 processed tabular features with 52 hand-engineered image features: channel statistics, center-vs-border color, symmetry, grayscale/texture proxies, and coarse RGB histograms.
Used patient ID as the cross-validation group so the same patient could not leak across folds, balanced training weights, and searched 120 HistGradientBoosting configurations. Selection considered both overall AUROC and the weakest age-group AUROC.
The best grouped-CV configuration reached a mean overall AUROC of 0.923 and a mean minimum age-group AUROC of 0.794. The project demonstrates multimodal feature engineering, leakage-aware evaluation, imbalance handling, and subgroup-aware model selection.
Course project. Source kept private; dataset attribution and assignment materials are not redistributed here.
Additional systems work
The portfolio leads with the projects above, but the underlying coursework also spans virtual machines, compression, assembly, interpreters, cache locality, and binary data representation.