Case study

Microplastic Detection

A full hardware-and-software pipeline for flagging microplastics in water. A Raspberry Pi acts as an edge gateway that reads an AS7265x 18-channel spectrometer and an MPU6050 inertial sensor; a laptop client polls it, scales the 18-band reading with a saved scaler, and classifies the sample with a PyTorch neural network, streaming the result to a live dashboard.

Overview and objective

Confirming microplastic contamination normally means laboratory instrumentation. The objective here was to build an end-to-end prototype that captures a spectral fingerprint near the sample, classifies it on the spot, and presents the result to an operator in real time - covering sensing, embedded control, model inference, and a live interface in one system.

The repository holds two phases. 01_SIH1_Old is the Phase 1 prototype (a heavier quantum-classical hybrid built on TensorFlow and PennyLane) and is archived. 02_SIH2_New is the current implementation: a conventional sensor stack and a PyTorch classifier.

Implemented

The current pipeline - sensor capture, the Flask gateway, the SpectroNet classifier, and the SocketIO dashboard - is implemented and runnable. The legacy Phase 1 code is retained for reference only.

System architecture

The design splits responsibilities across two machines: the Pi owns the hardware and a small HTTP API, while the laptop owns the model and the user interface.

Component map of the microplastic detection pipeline A four-stage stack: Pi sensors feed an edge gateway Flask API, which is polled by a laptop inference client, whose predictions feed a live dashboard. Sensing (Raspberry Pi) AS7265x 18-channel spectrometer (I2C) MPU6050 IMU and Pi camera Edge gateway (Pi, port 5001) Flask: /api/spectrometer, /api/imu, /video_feed, /action/<cmd> Inference (laptop, port 5002) StandardScaler + SpectroNet MLP 18 input features, 6 output classes Dashboard Flask + SocketIO telemetry Chart.js spectrum, Leaflet map, controls
Component map. The Pi also drives an ESP32-based drivetrain over serial. The gateway exposes read endpoints plus a command endpoint; the laptop client performs the model inference and serves the operator dashboard.

How data moves through the system

  1. Spectral and inertial capture

    The AS7265x returns 18 spectral bands over I2C, sampled and averaged to smooth out noise; the MPU6050 supplies orientation data used for stability handling.

    Implemented
  2. Edge serving on the Pi

    A Flask gateway exposes /api/spectrometer, /api/imu, a camera /video_feed, and an /action/<cmd> endpoint that forwards movement commands to the drivetrain over serial.

    Implemented
  3. Polling and scaling

    The laptop client polls the Pi at roughly 1 Hz and applies the saved scaler.pkl to normalise the 18-dimensional reading before inference.

    Implemented
  4. SpectroNet classification

    A PyTorch multi-layer perceptron produces a softmax over six classes - clean water, chlorophyll, and four polymers (PE, PP, PET, PS) - using weights from spectro_model.pth.

    Implemented
  5. Live telemetry to the dashboard

    The client emits updates over SocketIO; the dashboard renders the 18-band spectrum with Chart.js, plots position on a Leaflet map, and shows the current prediction.

    Implemented
  6. Manual control with a dead-man switch

    Drive commands are issued from the interface, and releasing the key stops the drivetrain so a dropped connection does not leave it moving.

    Implemented

Verified technology stack

  • ModelPyTorch SpectroNet MLP, 18 inputs to 6 classes, with ReLU, batch norm, and dropout layers (model.py)
  • Trained artifactsspectro_model.pth (weights) and scaler.pkl (feature scaler)
  • Web layerFlask with Flask-SocketIO; Chart.js and Leaflet on the front end
  • Hardware busesI2C (smbus2) for the AS7265x and MPU6050; serial (pyserial) for the ESP32 drivetrain
  • ImagingRaspberry Pi camera streaming MJPEG through the gateway
  • Configurationpython-dotenv (RPI_HOST, RPI_SERIAL_PORT, SECRET_KEY)
  • Training dataA generated synthetic spectra dataset (synthetic_spectro_data.csv), excluded from version control because of size

Key technical decisions and trade-offs

Split the edge from the inference host

The Pi owns the sensors and a small API; the laptop owns the model and the UI. This keeps the heavy PyTorch workload off the constrained edge device, at the cost of a network hop and a second process to run.

Polling at 1 Hz for simplicity

The client pulls samples on a fixed interval rather than the Pi pushing them. That is easy to reason about and recover, but the repository notes it adds latency compared with a push-based design.

Run even when the Pi is offline

The client starts with fallback data when the Raspberry Pi cannot be reached, so the interface and inference path can be developed and demonstrated without the full hardware rig present.

Harden flaky hardware reads

I2C and serial access use timeouts and retries so a dropped sensor read does not stall the pipeline, and drive commands are paired with a dead-man switch for safety.

Challenges and limitations

  • Network assumption. The system assumes a static internal network where the Pi address is known in advance.
  • Lighting sensitivity. Ambient lighting heavily affects AS7265x readings.
  • Polling latency. One-hertz polling is slower than a fully push-based telemetry design.
  • Synthetic training data. The classifier is trained on generated spectra; validation against real-world water samples is not demonstrated in the repository.
  • Fixed class set. The model distinguishes six configured classes and nothing outside them.
  • No reported accuracy. The evaluation script prints results but no accuracy metric is stored, so none is claimed here.

Reproducing and exploring it

Install dependencies once with pip install -r requirements.txt. Run the laptop client (it can start without the Pi, using fallback data):

cd 02_SIH2_New/02_Software/LaptopClient
python app_laptop.py

On the Raspberry Pi, deploy the hardware code and start the gateway:

cd 02_SIH2_New/01_Hardware/PI
python app.py

Point the client at the Pi with a .env file, and retrain the classifier with the dataset generator and training script if you want to reproduce the model from scratch.

Future engineering work

Future work - not implemented
  • Replace Pi-to-laptop polling with a pure WebSocket or MQTT push model.
  • Optimise SpectroNet for direct on-edge inference on the Raspberry Pi.
  • Wire in the GPS module logic that is currently only templated in the application.
  • Train and validate on a real-world capture set rather than synthetic spectra.

This project shares its edge-sensing and applied-model character with the QBioForge forensics build and the traffic-control system.