Design
Sentiment benchmarks usually report one aggregate number, but review language differs sharply by domain: electronics reviews are technical, fashion reviews are emotional. A model scoring 90% overall can be failing badly inside one category, and a business acting on its outputs would never know. The method: stratify the evaluation by product category, then run topic modelling on the reviews the models got wrong.
Execution as a Team of Five
The project ran in two graded stages: a ~3,000-word proposal (research gap, literature from lexicon methods to RoBERTa, methodology), then full execution as the module's final report and recorded presentation. Eighteen tasks were planned, owned and tracked across the team, from dataset ingestion and VADER baselines through BERT fine-tuning to the final write-up.
My Contribution
Across the two stages I worked on the research-gap framing and methodology, the data preparation pipeline, the comparative model evaluation with category-level breakdowns, the topic-modelling analysis of negative reviews, and co-presented the recorded final presentation.