The problem
Creative teams can reach powerful generative models, but individual provider interfaces do not create an operating system for production. Source files live in different places, model families expose different inputs, jobs can run for minutes, output links may expire, and experiments lose their connection to a client project or final deliverable.
ChuckOS makes the operational state around inference explicit. Users work inside projects, choose a focused workflow, select existing assets, configure a model, and receive the result back in the same library. The application owns identity, permissions, validation, storage, status, retry behavior, costs, and traceability while Replicate performs the GPU-intensive generation.
My role and scope
I owned the technical implementation end to end: translating the product concept into user flows, defining the domain model and API, building the React application and FastAPI service, integrating Replicate and AWS, and creating the deployment infrastructure. The visual identity and workflow concepts were product inputs; my responsibility was to turn them into a coherent, deployable system with clear boundaries between product logic, storage, and external inference.
How the product works
- 01
A user signs in, creates or selects a project, and uploads source images or videos.
- 02
ChuckOS validates the file, extracts media metadata, creates a thumbnail, and stores the original in a private S3 bucket.
- 03
The user chooses a workflow. The frontend builds its controls from JSON Schema and UI configuration stored with the selected model.
- 04
The API validates parameters, resolves asset UUIDs to time-limited URLs, records a durable job, and creates a placeholder output asset.
- 05
Replicate runs the prediction while the browser stays responsive and follows lightweight status updates.
- 06
An idempotent webhook reconciles completion, and ChuckOS copies successful provider output into its own S3 namespace.
- 07
The result becomes a first-class project asset that can be inspected, tagged, favourited, regenerated, downloaded, or reused.

Architecture and implementation
Frontend generated from model metadata
The React 19 and TypeScript client uses React Router, Vite, Tailwind, Zustand, Axios, and AJV. Instead of hard-coding a form for every model, a reusable engine turns Draft 7 JSON Schema plus defaults, grouping, field order, widget hints, and help text into the correct controls. Dedicated workflows can promote a few important inputs without breaking the shared contract.
Backend orchestration around inference
FastAPI owns authentication, authorization, validation, business rules, and provider integration. SQLAlchemy models represent users, projects, media assets, tags, model configuration, system settings, and jobs; Alembic manages the schema. Focused services separate uploads, media processing, S3 access, job orchestration, and model administration from HTTP routing.
Stable media identity and lifecycle
PostgreSQL is the source of truth for structured state. Images and videos share a polymorphic asset model and stable public UUIDs, while binary originals and thumbnails live in S3 behind presigned URLs. Upload orchestration cleans up partial state if processing fails, and generated output inherits the same project, filtering, tagging, and download behavior as an uploaded file.
Infrastructure as part of the product
Terraform defines the AWS deployment in eu-west-1: an ALB with managed TLS, separate frontend and backend containers on EC2, private RDS PostgreSQL, a private versioned S3 bucket, ECR, IAM, networking, and a budget alert. GitHub Actions runs tests and type/build checks before publishing commit-tagged container images.

Key decisions and trade-offs
Configuration-driven model onboarding
One form and validation pipeline supports multiple model families. The trade-off is that every new inference host still needs an execution adapter and callback strategy.
Durable jobs before a queue
Persisting a job and placeholder asset before provider submission keeps the MVP small and recoverable. Higher throughput would justify workers and an outbox.
Owning generated output
Copying successful results out of temporary provider storage preserves the product lifecycle. Large video volumes would move ingestion to dedicated workers.
Compact cloud footprint
A single application node keeps the MVP understandable and cost-conscious. Autoscaling, managed secrets, and multi-node availability remain later-stage improvements.

A complete path from source asset to organized output
ChuckOS delivers six focused workflow entry points, a searchable project library, configurable model administration, and user/admin reporting. The technical review recorded all 26 backend tests passing, a successful TypeScript check, and a clean frontend production build.
The project demonstrates the engineering that surrounds an AI call: a reliable asset lifecycle, asynchronous state, schema-driven product controls, cost visibility, and a repeatable deployment path.
Reflection
The configuration-driven model layer and first-class asset lifecycle were the right foundations. The next stage would formalize a provider adapter interface, introduce background workers and an outbox for media ingestion, harden browser sessions, and split the single compute node into independently scalable services. Those changes extend the current domain model rather than replacing it.