MIS · Santa Clara University

Hieu "Ray" Bui

Turning data into practical insight.

Scroll

Background

About Me

Hieu Ray Bui

I'm a Management Information Systems student at Santa Clara University with a focus on data analytics, machine learning, and business-oriented problem solving. I enjoy work that sits at the intersection of technical depth and real-world impact.

Through coursework and hands-on projects, I've built end-to-end machine learning pipelines, designed relational databases, and led a team through a 24-hour AI hackathon. I'm currently exploring early-career opportunities in data analytics and data science.

Outside of data, I care about clear communication — knowing how to explain what a model does and why it matters is just as important as building it.

Languages & Tools

  • Python (Pandas, Scikit-Learn)
  • SQL & Database Design
  • Matplotlib / XGBoost
  • Git & GitHub

Cloud & Other

  • AWS (EC2, S3, Bedrock)
  • Docker & Linux
  • Jupyter Notebooks
  • Data Visualization

Research

Predicting Antidepressant Prescribing Patterns with Open-Source LLMs

Research Assistant · OCHIN LLM Evaluation Project · Advised by Prof. Jingbo Hou, Santa Clara University

Can a language model predict how doctors actually prescribe medication, just from a patient's basic demographics and diagnosis? I've been testing this question against real clinical data from OCHIN, a network of community health centers, across two benchmark tasks covering 896+ patient demographic groups, three open-source models spanning different architectures and parameter scales, and three tiers of prompt guidance.

The core finding: even with detailed clinical context, LLMs consistently fall short of a naive baseline that simply predicts the population average. Traditional machine learning models, trained directly on the real data, do better. But the gap isn't about clinical reasoning; the models correctly identify which drugs are generally more or less common. It's about grounding: without access to the actual dataset, general medical knowledge alone isn't enough to predict fine-grained, group-specific patterns.

Along the way, I built and debugged the full inference pipeline from scratch on SCU's WAVE HPC cluster, running open-source models through Ollama, and caught a subtle class-mapping bug that was silently corrupting results before it reached evaluation.

This work is ongoing — the next phase expands testing to an 18-model roster spanning different architectures, sizes, and reasoning capabilities, and moves toward individual patient-level prediction.

Python SLURM Ollama LLM Evaluation HPC

Get In Touch

Let's Connect

I'm actively exploring internship and early-career opportunities in data analytics and data science.