Computer Science + Mathematics @ Tufts University

Joel Lawore

I build and optimize systems, data pipelines, and machine-learning models — especially where performance, evaluation, and real-world constraints matter.

joellawore@gmail.com Recruiting for 2027 full-time SWE / FDE roles

Selected work

Projects with real technical signal.

Chosen for depth, engineering decisions, measurable outcomes, and breadth across systems, ML, and product work.

Python NLP scikit-learn

Readability Classification

Built a text-classification pipeline combining character 3–5-gram TF-IDF with 28 readability and stylometric features.

Key signal: caught optimistic validation caused by same-book passages crossing folds, then moved to author/title-grouped CV.
Read case study
Python Image features Gradient boosting

Skin-Lesion Classification

Combined clinical metadata with engineered image features, using patient-grouped evaluation and class weighting on a heavily imbalanced dataset.

107total features
92.4%positive rate
Read case study
C++ Hash tables Filesystem traversal

Gerp — File Search Utility

Indexed text across recursive directory trees with a custom templated hash table using linear probing and rehashing, then exposed case-sensitive and case-insensitive CLI search.

Source kept private in accordance with course policy.

CSS Stylus UI systems

Tufts SIS Modern

Designed a modular, reversible visual modernization layer for Tufts’ legacy PeopleSoft SIS while preserving existing behavior and navigation.

Architecture: core design tokens plus separate navigation, dashboard, table, form, dark-mode, and experimental modules.

Case studies

How the engineering decisions actually worked.

These are intentionally implementation-focused without publishing restricted course solutions.

01 / Programming languages + performance

uScheme Interpreter Optimization

3.8×faster on the assignment benchmark

Problem

The interpreter created closures by carrying the surrounding environment. That is semantically safe, but potentially wasteful when a lambda references only a small subset of those bindings.

What changed

Implemented recursive free-variable analysis over the expression tree, including binding forms, nested lambdas, mutation, control flow, and function application. Closure creation then filtered the captured environment to bindings proven relevant to the lambda.

Engineering process

Added targeted checks around free-variable behavior, integrated the analysis at closure creation, and benchmarked the modified interpreter against the original source implementation.

Result + signal

Runtime fell from 5.14 seconds to 1.35 seconds on the assignment benchmark — a ~74% reduction. The project demonstrates interpreter internals, lexical scope, recursive AST analysis, and performance optimization without changing language semantics.

Source kept private in accordance with course policy.

02 / Systems + performance engineering

Universal Machine Performance Optimization

C + Callgrindprofiling-guided optimization

Problem

Start with an unfamiliar Universal Machine implementation and make it faster without changing its instruction semantics. The initial design relied on higher-overhead container abstractions for frequently accessed runtime state.

Technical decisions

Progressively moved segment and reusable-ID representations from Hanson sequences to UArrays and then flatter arrays, and consolidated execution paths to reduce indirection and improve locality.

Debugging + validation

A consolidation pass unexpectedly exhausted stack space after moving state into local scope; the fix was to allocate the large runtime state on the heap. Callgrind identified instruction dispatch as the dominant hot path, and generated assembly confirmed the compiler lowered the opcode switch to a jump table with frequently used state kept largely in registers.

What it demonstrates

Performance work as an evidence loop: profile first, change representation based on a bottleneck hypothesis, debug secondary effects, then inspect machine-level behavior rather than assuming the compiler did what the source suggested.

Team project. Source kept private in accordance with course policy.

03 / ML evaluation + NLP

Readability Classification

28 + TF-IDFengineered + text features

Problem

Classify literary passages into two reading-level groups from text plus computed readability and stylometric metadata.

Model

The final pipeline used character-level 3–5-gram TF-IDF alongside 28 numerical features, standardized the numeric branch, and trained logistic regression.

The important discovery

Ordinary 10-fold validation produced a mean AUROC around 0.82, but inspection showed that passages from the same book could land in both training and validation folds. That made the model partly answer “have I seen this work before?” instead of the intended reading-level question.

Evaluation redesign

Rebuilt validation around groups defined by author + title using stratified group folds, so no literary work crossed the train/validation boundary. The lower, more variable grouped scores were treated as more trustworthy than the prettier leaked result.

Course project. Source kept private; no assignment solution code is published here.

04 / Applied ML + multimodal features

Skin-Lesion Classification

0.923best mean grouped-CV overall AUROC

Problem

Predict a binary skin-lesion target from patient metadata plus smartphone images. The training data contained 1,178 examples and a severe 92.4% positive-class imbalance.

Feature design

Combined 55 processed tabular features with 52 hand-engineered image features: channel statistics, center-vs-border color, symmetry, grayscale/texture proxies, and coarse RGB histograms.

Evaluation + model selection

Used patient ID as the cross-validation group so the same patient could not leak across folds, balanced training weights, and searched 120 HistGradientBoosting configurations. Selection considered both overall AUROC and the weakest age-group AUROC.

Result + signal

The best grouped-CV configuration reached a mean overall AUROC of 0.923 and a mean minimum age-group AUROC of 0.794. The project demonstrates multimodal feature engineering, leakage-aware evaluation, imbalance handling, and subgroup-aware model selection.

Course project. Source kept private; dataset attribution and assignment materials are not redistributed here.

Additional systems work

Broader low-level foundation.

The portfolio leads with the projects above, but the underlying coursework also spans virtual machines, compression, assembly, interpreters, cache locality, and binary data representation.

32-bit Universal Machine in C PPM image compression + bit packing UM assembly RPN calculator Huffman compression in C++ Typed RPN interpreter in C++ Cache-locality experiments

Contact

Interested in the work?

I’m recruiting for 2027 full-time software engineering, forward-deployed engineering, and related technical roles.