Ben Collyer

DevOps Engineer · Platform Engineer · SRE

Open to remote West Midlands, UK linkedin.com/in/ben-c-58634874

DevOps and Platform Engineer with 13+ years across Linux, AWS and infrastructure automation. I build and run the systems businesses depend on, and I write the application code that runs on them. Strongest on Linux and AWS infrastructure, CI/CD pipelines and Python automation.

Selected work

Running a commercial platform single-handed

Sole engineer for buylist.co.uk since 2023, infrastructure and application both. The client resold it to a second card business as their operating software, and I re-architected it for multi-tenancy: one codebase, one pipeline, two independent shops.

Disaster recovery, then the real thing

Designed and delivered AWS Elastic Disaster Recovery across Jadu's data centre estate. When a major outage hit, I led the failover to EC2: critical business operations restored overnight, 75% of services recovered within 48 hours.

RAG in production, not in a notebook

Built and shipped the backend for Agentex, Jadu's conversational search product: Firecrawl, OpenAI embeddings, PGVector and GPT with guardrails, deployed via ArgoCD. Agentex was a key part of Jadu's acquisition by Netcall in 2025.

Three and a half years on the rota

Held the production on-call rota at Jadu for the full time I was there, handling live incidents across the customer estate. Reliability as a standing commitment rather than a project.

Experience

Oct 2023→ present

Freelance Platform & Full Stack Engineer, buylist.co.uk

Part-time, alongside full-time employment · Trading-card buyback and retail business

  • Sole engineer for the platform end to end, infrastructure as well as application: Linux hosts, Caddy and gunicorn under supervisor, MariaDB and Redis, a separate UAT environment, and deployment through GitLab CI on a self-hosted runner. Live and load-bearing for three years, with no second pair of hands.
  • The client resold the platform to a second card business as their operating software, which I then re-architected for multi-tenancy: one codebase and one pipeline, two independent shops, separate databases and branding. Within 18 months its automations had freed enough staff time from manual pricing and listing to support the shop's move into larger premises with a storefront and event space.
  • Scheduled job pipeline handling overnight repricing, inventory recounts, marketplace order reconciliation every few minutes and Meilisearch reindexing over a 112,000-card catalogue, against a price-history table of 11.6M rows.
  • Two-way CardMarket integration for pricing, order import and fulfilment, plus a bulk listing pipeline ingesting scanner output for thousands of cards: reprice, reconcile against stock and publish in one pass. Two-way Shopify sync with webhook callbacks, and a published Firefox extension injecting platform data into CardMarket's own pages for staff.
  • Pickstream, the picking interface, built around how the warehouse physically works rather than around the database: tablets on stands mounted above the picking stations, driven by a finger-worn scroll wheel, so staff advance through a pick at speed without ever taking a hand off the stock or touching a screen. Designed from repeated visits to the client's floor, watching how pickers actually move.
  • Wider warehouse operations tooling over a Flask and MariaDB backend of ~24,000 lines across 37 models: holding-cell allocation, label printing and per-set stock valuation, with role-based access, TOTP 2FA and an admin audit trail.

Python · Flask · SQLAlchemy · MariaDB · Meilisearch · SvelteKit · Tailwind · Docker · GitLab CI · Caddy · Linode

Nov 2022→ Jul 2026

DevOps Engineer, Jadu / Netcall

Remote · Digital experience and CRM software for UK public sector

  • Designed and delivered AWS Elastic Disaster Recovery (DRS) across the company's data centre estate. During a major outage this enabled a full failover to EC2: critical business operations restored overnight and 75% of services recovered within 48 hours.
  • Built and shipped the backend for Agentex, Jadu's AI-powered conversational search product: a full RAG pipeline using Firecrawl, OpenAI embeddings, PGVector and GPT for structured response generation with guardrails, deployed via GitOps with ArgoCD. Agentex was a key part of Jadu's acquisition by Netcall in 2025.
  • Held the production on-call rota for the full 3.5 years, handling live incidents across the customer estate.
  • Provisioned and managed customer AWS environments entirely as code: Terraform applied through GitLab CI, with AWX orchestrating Ansible for configuration. Progressively migrated the remaining legacy data centre servers to AWS with modernised, refined server roles.

AWS (EC2, RDS, ElastiCache, IAM, WAF) · Terraform · Ansible · AWX · GitLab CI · ArgoCD · Docker · Docker Compose · Python · PostgreSQL · MariaDB · Rocky Linux · CentOS · Ubuntu

Aug 2022→ Nov 2022

Systems Engineer, Spectrum.LIFE

Remote · Digital mental health and wellbeing platform

  • Containerised the core Node.js/PHP application with Docker, improving deployment consistency.
  • Managed AWS infrastructure: EC2, RDS, CloudFront, Route 53, VPCs and load balancers.
Nov 2021→ Aug 2022

DevOps Engineer, T-Shirt & Sons

Remote · High-volume print-on-demand and screen-print apparel manufacturing and fulfilment

  • Managed and extended GitLab CI/CD infrastructure across Docker, Docker Compose and Docker Swarm.
  • Built and self-hosted internal Python (Flask) tooling on a highly available Proxmox cluster I built and maintained end to end.
  • Owned Traefik configuration and reverse proxy routing across internal services.

GitLab CI/CD · Docker Swarm · Traefik · Proxmox · Python · Flask · Azure DevOps

Mar 2020→ Nov 2021

DevOps / Full Stack Developer, Chrysalis Loyalty

Remote · Automotive customer loyalty and retention software; acquired by Autofutura in 2021

  • Managed CI/CD pipelines, branching strategies and software releases, and maintained cloud infrastructure on AWS (EC2, S3, RDS) and Azure DevOps.
  • Wrote Bash and Python automation for VM management, service reliability and data migrations.
  • Contributed Laravel and Flask development on Key2Key, the company's cloud lead-generation platform.

AWS · Azure DevOps · Bash · Python · Laravel · Flask · ElasticSearch · Cassandra · Redis · MariaDB

Mar 2019→ Nov 2019

Full Stack Developer / SRE, XQ Cyber Resilience

Cybersecurity software and services

  • Combined application development with site reliability work, maintaining service reliability and infrastructure stability across the product estate.
Aug 2018→ Mar 2019

Full Stack Developer, Online Home Retail (Plumbworld)

One of the UK's largest online-only bathroom and kitchen retailers, trading as Plumbworld

  • Web application development and maintenance.
Sep 2012→ Apr 2017

System Administrator, Matcon Ltd / IDEX Corp.

Global industrial equipment manufacturer, part of IDEX Corporation

  • Maintained internal IT infrastructure, Windows and Linux servers, networking and end-user systems across the business.

Personal projects & homelab

Shippedcodebeasts.net

CodeBeasts

An Android creature-collector designed and shipped solo, currently in Google Play closed testing ahead of production release.

  • Beasts, names, sprites and stats are derived deterministically from hashed barcodes, so the game needs no asset packs and no server-side state.
  • Offline peer-to-peer trading and battles over Google Nearby Connections.
  • Cloud save and restore through Play Games, using a portable HMAC-signed save format so collections survive a device change.
  • Health Connect integration converting step counts into in-game progression.
  • Versioned, signed and R8-minified release builds shipped through Play Console tracks, backed by a JVM test suite and a static SvelteKit site.

Kotlin · Jetpack Compose · Gradle · Play Console · Nearby Connections · Health Connect · SvelteKit

Runningself-hosted

Homelab

Production-shaped infrastructure, used to work with technologies ahead of adopting them professionally.

  • Kubernetes cluster: multi-node cluster running self-hosted services, managed day to day with k9s.
  • GitOps: cluster state declared in Git and reconciled with ArgoCD, mirroring the workflow used in production at Jadu.
  • Storage: Longhorn for distributed block storage and replicated persistent volumes.
  • Virtualisation: Proxmox cluster hosting the underlying nodes.
  • Local AI: llama.cpp and Ollama for local model inference; experimentation with Hugging Face models, RAG pipelines and prompt engineering.

Kubernetes · ArgoCD · Longhorn · Helm · k9s · Proxmox · Docker · llama.cpp · Ollama

Skills

Infrastructure as code & automation

Terraform, Ansible, AWX, Bash, Python

Containers & orchestration

Kubernetes, Docker, Docker Compose, Docker Swarm, Helm, Longhorn, k9s

CI/CD & GitOps

GitLab CI/CD, Azure DevOps, ArgoCD, release management, branching strategies

Cloud & infrastructure

AWS (EC2, RDS, ElastiCache, S3, IAM, WAF, CloudFront, Route 53, VPC, Load Balancers, Elastic Disaster Recovery), Proxmox, Traefik, disaster recovery, high availability

Operating systems

Linux, Rocky Linux, CentOS, Ubuntu, Windows Server

Databases

PostgreSQL, MariaDB, MySQL, Redis, ElasticSearch, Cassandra

Development & scripting

Python, Flask, FastAPI, Bash, PHP, Laravel, Node.js

AI & ML engineering

RAG pipelines, OpenAI embeddings, PGVector, Firecrawl, Crawl4AI, llama.cpp, Ollama, Hugging Face, LLM integration, prompt engineering, guardrails

Education & professional development

Self-taught engineer, 13+ years of production experience across system administration, software development and infrastructure engineering.

2017 to 2018 · Full-time self-directed study. Left system administration to retrain intensively in Linux and web development. Returned to full-time development work in Aug 2018 and have worked in DevOps and platform roles since.

Ben Collyer · West Midlands, UK · open to fully remote roles