A full hardware-and-software pipeline for flagging microplastics in water. A Raspberry Pi acts as an edge gateway that reads an AS7265x 18-channel spectrometer and an MPU6050 inertial sensor; a laptop client polls it, scales the 18-band reading with a saved scaler, and classifies the sample with a PyTorch neural network, streaming the result to a live dashboard.
Confirming microplastic contamination normally means laboratory instrumentation. The objective here was to build an end-to-end prototype that captures a spectral fingerprint near the sample, classifies it on the spot, and presents the result to an operator in real time - covering sensing, embedded control, model inference, and a live interface in one system.
The repository holds two phases. 01_SIH1_Old is the Phase 1 prototype (a heavier quantum-classical hybrid built on TensorFlow and PennyLane) and is archived. 02_SIH2_New is the current implementation: a conventional sensor stack and a PyTorch classifier.
Implemented
The current pipeline - sensor capture, the Flask gateway, the SpectroNet classifier, and the SocketIO dashboard - is implemented and runnable. The legacy Phase 1 code is retained for reference only.
System architecture
The design splits responsibilities across two machines: the Pi owns the hardware and a small HTTP API, while the laptop owns the model and the user interface.
Component map. The Pi also drives an ESP32-based drivetrain over serial. The gateway exposes read endpoints plus a command endpoint; the laptop client performs the model inference and serves the operator dashboard.
How data moves through the system
Spectral and inertial capture
The AS7265x returns 18 spectral bands over I2C, sampled and averaged to smooth out noise; the MPU6050 supplies orientation data used for stability handling.
Implemented
Edge serving on the Pi
A Flask gateway exposes /api/spectrometer, /api/imu, a camera /video_feed, and an /action/<cmd> endpoint that forwards movement commands to the drivetrain over serial.
Implemented
Polling and scaling
The laptop client polls the Pi at roughly 1 Hz and applies the saved scaler.pkl to normalise the 18-dimensional reading before inference.
Implemented
SpectroNet classification
A PyTorch multi-layer perceptron produces a softmax over six classes - clean water, chlorophyll, and four polymers (PE, PP, PET, PS) - using weights from spectro_model.pth.
Implemented
Live telemetry to the dashboard
The client emits updates over SocketIO; the dashboard renders the 18-band spectrum with Chart.js, plots position on a Leaflet map, and shows the current prediction.
Implemented
Manual control with a dead-man switch
Drive commands are issued from the interface, and releasing the key stops the drivetrain so a dropped connection does not leave it moving.
Implemented
Verified technology stack
ModelPyTorch SpectroNet MLP, 18 inputs to 6 classes, with ReLU, batch norm, and dropout layers (model.py)
Trained artifactsspectro_model.pth (weights) and scaler.pkl (feature scaler)
Web layerFlask with Flask-SocketIO; Chart.js and Leaflet on the front end
Hardware busesI2C (smbus2) for the AS7265x and MPU6050; serial (pyserial) for the ESP32 drivetrain
ImagingRaspberry Pi camera streaming MJPEG through the gateway
Training dataA generated synthetic spectra dataset (synthetic_spectro_data.csv), excluded from version control because of size
Key technical decisions and trade-offs
Split the edge from the inference host
The Pi owns the sensors and a small API; the laptop owns the model and the UI. This keeps the heavy PyTorch workload off the constrained edge device, at the cost of a network hop and a second process to run.
Polling at 1 Hz for simplicity
The client pulls samples on a fixed interval rather than the Pi pushing them. That is easy to reason about and recover, but the repository notes it adds latency compared with a push-based design.
Run even when the Pi is offline
The client starts with fallback data when the Raspberry Pi cannot be reached, so the interface and inference path can be developed and demonstrated without the full hardware rig present.
Harden flaky hardware reads
I2C and serial access use timeouts and retries so a dropped sensor read does not stall the pipeline, and drive commands are paired with a dead-man switch for safety.
Challenges and limitations
Network assumption. The system assumes a static internal network where the Pi address is known in advance.
Polling latency. One-hertz polling is slower than a fully push-based telemetry design.
Synthetic training data. The classifier is trained on generated spectra; validation against real-world water samples is not demonstrated in the repository.
Fixed class set. The model distinguishes six configured classes and nothing outside them.
No reported accuracy. The evaluation script prints results but no accuracy metric is stored, so none is claimed here.
Reproducing and exploring it
Install dependencies once with pip install -r requirements.txt. Run the laptop client (it can start without the Pi, using fallback data):
cd 02_SIH2_New/02_Software/LaptopClient
python app_laptop.py
On the Raspberry Pi, deploy the hardware code and start the gateway:
cd 02_SIH2_New/01_Hardware/PI
python app.py
Point the client at the Pi with a .env file, and retrain the classifier with the dataset generator and training script if you want to reproduce the model from scratch.
Future engineering work
Future work - not implemented
Replace Pi-to-laptop polling with a pure WebSocket or MQTT push model.
Optimise SpectroNet for direct on-edge inference on the Raspberry Pi.
Wire in the GPS module logic that is currently only templated in the application.
Train and validate on a real-world capture set rather than synthetic spectra.
Related case studies
This project shares its edge-sensing and applied-model character with the QBioForge forensics build and the traffic-control system.