Research

I am an applied machine learning researcher and Computer Science PhD candidate at Binghamton University (SUNY), working in the RT Lab under Prof. Kyoung-Don Kang. My research focuses on multimodal retrieval and vision-language models, with an emphasis on making frozen models more accurate, robust, and efficient without costly retraining. As an Applied Research Intern on eBay's Ads team, I fine-tuned BLIP with visual-grounding late interaction for fashion product retrieval, developed LoCaR, a training-free framework for localized retrieval with frozen backbones, and built LOCFG, a leakage-controlled benchmark for small-referent retrieval.

My broader work covers reliability under distribution shift: post-hoc out-of-distribution detection with LSF, efficient test-time adaptation (Sensors), and robust edge AI for insect pest monitoring over LoRa networks (ICDCN 2027). I have also worked on security in source-free domain adaptation (ICCV 2023) and unsupervised learning (IJCNN 2021), and previously built long-context summarization models as a Research Scientist at Cobalt Speech & Language.


LoCaR: Localized Canvas Retrieval with Frozen Vision-Language Models

LoCaR region-canvas retrieval pipeline
Whole-image embeddings can miss the small objects and local attributes that localized text queries describe. LoCaR is a training-free framework that turns each SAM2 proposal into a region canvas (a crop with a blurred background), encodes it independently with a frozen vision-language model, and fuses the best region match with whole-image similarity. Across CLIP, SigLIP2, and Qwen3-VL-Embedding, it outperforms TextRegion on RefCOCOg. We also introduce LOCFG, a small-referent retrieval benchmark derived from Visual Genome. This work was carried out during my Applied Research internship at eBay.

LSF: Layer Selection and Fusion for Post-Hoc OOD Detection

LSF layer selection and fusion framework
Distance-based OOD detectors usually rely only on the penultimate layer, which may suppress lower-level variation that helps reveal distribution shifts. LSF picks the intermediate layer with the largest entropy-density drop and concatenates it with the penultimate representation, then scores inputs with a shrinkage-regularized Mahalanobis distance. It needs no OOD data or fine-tuning, handles unequal layer widths, and improves near-OOD, far-OOD, and covariate-shift detection across ViT, Swin, and ConvNeXt backbones.

Efficient Test Time Adaptation

TDA-L architecture diagram
Adapting vision–language models to test-time distribution shifts, such as lighting or weather changes, usually relies on backpropagation, which is too slow for real-time sensing. Building on the training-free, cache-based TDA adapter, TDA-L applies pre-learned Low-Rank Adaptation (LoRA) matrices to query and cached features to shrink their size at inference. Across seven benchmarks, TDA-L maintains accuracy while lowering latency and memory use and increasing throughput. Published in Sensors.

A DNN for Detecting Spotted Lanternflies Using Energy Efficient WAN

Spotted Lanternfly detection system
For my Master's thesis, I built an energy-efficient system for detecting Spotted Lanternflies (SLF), an invasive pest that is hard to monitor in remote spots such as tree branches. A quantized MobileNetV3 model trained on a dedicated SLF dataset runs on a Raspberry Pi with a LoRa module, achieving high accuracy and low latency on the device. Compared with WiFi and Bluetooth, the LoRa-based system covers a larger area while using significantly less power.

WASP: Edge-AI Insect Pest Monitoring in Challenging Environments

RPCCL three-stage curriculum overview
Field-deployed pest monitors face two coupled problems: lightweight classifiers lose accuracy under rain, fog, glare, and low light, and LoRaWAN links become unreliable in dense vegetation. WASP addresses both. RPCCL (Randomized Progressive Curriculum Contrastive Learning) trains lightweight models through a three-stage curriculum that increases corruption severity and diversity, with random sampling within each stage, at no extra inference cost. LA-ADR (Location-Aware Adaptive Data Rate) uses line-of-sight, vegetation density, and clustering to assign LoRa data rates. Across benchmark corruptions, ns-3 simulations, and a 10-node forest deployment, WASP improves both classification robustness and packet reception. Accepted at ICDCN 2027.

A DNN for Detecting Spotted Lanternflies Using Energy Efficient WAN

Spotted Lanternfly detection system
For my Master's thesis, I built an energy-efficient system for detecting Spotted Lanternflies (SLF), an invasive pest that is hard to monitor in remote spots such as tree branches. A quantized MobileNetV3 model trained on a dedicated SLF dataset runs on a Raspberry Pi with a LoRa module, achieving high accuracy and low latency on the device. Compared with WiFi and Bluetooth, the LoRa-based system covers a larger area while using significantly less power.

Adaptive Real Time Object Detection using advance optical flow algorithm

Traffic scene used for real-time object detection
We are developing a real-time object detection system that can adapt to changes in the environment. We are using optical flow to detect changes in the environment, and then using this information to adapt the object detection model to the new environment.

Investigating Linear Neural Network’s Vulnerability

Comparison of neural network architectures
In this research, the primary objective is to investigate the effect of elimination of non linear activation in the DNN in terms of robustness. We have shown that the linear neural network is vulnerable to adversarial attacks.

Security Threat in Source Free Domain Adaptation

Triggerless backdoor attack illustration
We investigated the effect of a source adversary which may inject a hidden malicious behavior (Backdoor/Trojan) during source training and potentially transfer it to the target domain even after benign training by the victim (target do-main owner). We also built a defense method for the attack as well.

Weight Pruning

Weight pruning illustration
We investigated the effect of weight pruning in unsupervised learning setup. We also proposed a weight perturbation method. We showed that the weight pruning is very effective in unsupervised learning setup.

Local features detection using 3D point cloud

3D point cloud of a face
We built a pipeline to detect local features using 3D point cloud. The pipeline consists of Kmeans clustering algorithm to first segment the point cloud and then we used the Gaussian Mixture Model to get the local features. We then ran the pipeline on the 3D point cloud of the face and got the local features.