A food-scanning prototype that spans a mobile app, a web client, and a Python inference service. A user photographs a food item; a multi-task ResNet-50 classifies it, and the backend attaches basic nutritional and Ayurvedic (Dosha) reference information where the reference data supports it. It is deliberately framed as a prototype, not a clinical or nutritional authority.
The goal was to build one inference service and put it behind real clients - a camera-first mobile experience and a browser experience - rather than a single script. That forced the interesting engineering work: bounding memory on a CPU host, validating uploads, and returning a structured response that both clients can render.
Each scan flows through a single documented endpoint, POST /api/v1/scan/food, which returns the recognised food, a confidence value, optional reference information, and model metadata.
Stated limitation
Model accuracy for the 101-class problem is not yet verified, and the application is not a clinical or nutritional authority. The nutrition and Ayurveda fields are advisory reference data, not instructions.
System architecture
Three surfaces, one inference service. The clients stay thin; all recognition and mapping happens in the backend.
Component map. The three model heads share the ResNet-50 backbone; the food head drives the recognised label, while the auxiliary heads feed the reference mapping. Docker Compose runs the frontend on 8080 and the backend on 8000.
How data moves through the system
Capture
The mobile app captures a photo with Expo Camera; the web client accepts an uploaded image file.
Implemented
Upload and validation
The image is posted as multipart form data to /api/v1/scan/food. JPEG, PNG, and WebP are accepted up to 5 MiB; other content types and oversized files are rejected.
Implemented
CPU inference
The multi-task ResNet-50 runs on CPU. The model loads lazily on the first request under a lock, so startup stays quick and the memory spike is bounded.
Implemented
Reference mapping
The response builder attaches nutritional and Ayurvedic (Dosha) reference data from a static mapping file. When an item has no mapping, the field reports that information is unavailable instead of inventing values.
Implemented
Structured response
The API returns JSON containing the food label, a confidence value, the optional reference sections, and the model name and version.
Implemented
Client rendering
Both clients display the result. The mobile API client adds retries with backoff, request timeouts, and secure storage for its configuration.
Implemented
Verified technology stack
BackendFastAPI served by Uvicorn, with Pydantic schemas
ModelMulti-task PyTorch ResNet-50 (mtl_model_latest.pth, about 95 MB) with a Food-101 class list
ML librariesPyTorch and torchvision, pinned in the backend requirements
Web clientReact with Vite, plus Nginx in the container image
Mobile clientReact Native on Expo (Expo SDK 54), using expo-camera
Data and configSQLite via SQLModel (the scanner path is currently stateless); environment variables for origins and API URL
PackagingDocker Compose: web frontend on 8080, FastAPI backend on 8000
Key technical decisions and trade-offs
A single backend worker on purpose
The PyTorch model consumes substantial RAM. The backend is deliberately pinned to one Uvicorn worker to avoid out-of-memory failures on small hosts, trading concurrency for reliability.
Lazy model loading under a lock
The ~95 MB model is loaded on first use rather than at import. That keeps cold start fast and avoids duplicate loads under concurrent requests.
A mapping layer that admits gaps
Nutrition and Dosha data come from a small static mapping file. Unsupported foods return a "not available" note rather than fabricated advice.
CPU-only inference
Running on CPU avoids a GPU requirement and keeps the image portable; the trade-off is inference time of roughly 0.1 to 0.3 seconds per image after the initial load.
Challenges and limitations
Unverified accuracy. The 101-class model and its real-world behaviour are not yet validated, so no accuracy figure is claimed.
Not authoritative. Nutritional and Ayurvedic output is reference guidance only and must not be treated as medical or dietary advice.
Reference-data coverage. The mapping layer is a small static file, not an exhaustive database.
Not deployed. Cloud deployment is planned but not done; development origins are hardcoded in the compose file for local testing.
Resource footprint. The model is memory-hungry, which constrains where the backend can run.
Optional / experimental
The repository also contains separate experiment code for a quantum-circuit classifier. It is not part of the scan API or the primary runtime and is noted here only for completeness.
Reproducing and exploring it
The fastest path is Docker Compose, which runs the full stack locally:
docker compose build
docker compose up -d
docker compose ps
The web client is then on http://localhost:8080 and the API on http://localhost:8000. To run natively instead:
The web client starts with npm install and npm run dev in 02_Frontend; the mobile app starts with npm install and npx expo start in 03_Mobile, pointed at a reachable LAN address for a physical device.
Future engineering work
Future work - not implemented
Validate model accuracy on Food-101 and on real user photographs.
Broaden the nutrition and Ayurveda reference data, or integrate a proper source.
Deploy the stack with production environment configuration instead of local defaults.
Related case studies
Like the other builds here, SecretEye pairs an applied model with real clients; it shares its Python service style with the QBioForge and microplastic projects.