Skip to main content

My server pays for a voice AI call it never hears

· 11 min read
Pascal Nehlsen
Platform & Security Engineering

Falar is a speaking tutor for European Portuguese. You talk to Ana, a teacher from Lisbon, and she answers out loud, corrects you, and remembers which mistakes you keep making. It has been in internal testing on Google Play since 1 October.

The architecture decision that shapes everything else: the phone talks to OpenAI's Realtime API directly, over WebRTC. My backend creates the call and holds the API key, but the audio never passes through it. That keeps the round trip near one second, and it is the only way I found that lets the model actually hear how a learner pronounces a word rather than read a transcript of it.

It also means the most expensive thing in the system runs on a device I do not control, on a connection I cannot see, billed to my account.

Bound to 127.0.0.1 and still reachable from every browser tab

· 13 min read
Pascal Nehlsen
Platform & Security Engineering

On 5 October I made CaptureDesk public. It is a small Electron app: Loom has no desktop client for Linux, and recording in the browser gives you no camera bubble that floats over other windows, no drawing on screen and no global shortcuts. CaptureDesk wraps Loom's Record SDK and adds them.

Before flipping the repository to public, I did a hardening pass and found three real problems. One of the fixes was binding the app's local server to 127.0.0.1 instead of every interface.

The next day I read it again and found that the server still answered any web page open in the user's browser. Binding to loopback had not closed that door. It had only made it narrower.

An agent that writes production code needs a gate it cannot open

· 14 min read
Pascal Nehlsen
Platform & Security Engineering

I let Claude write workflows and deploy them to my n8n instance. Not suggestions I paste in by hand. Actual workflow.json, uploaded through a pipeline, running against real calendars, real mailboxes, real spreadsheets.

The thing I was afraid of was never bad code. Bad code fails loudly, usually on the first run, and you fix it.

How to red-team your own agent sandbox

· 11 min read
Pascal Nehlsen
Platform & Security Engineering

Last time I reported a result from the lab: the sandbox holds, and the model is what stops the leak. Inside an intact microVM, the egress allow-list contributed nothing when data left over a host it had to permit, and the only control that actually stopped a leak was the model's own alignment.

The fair question in response was: how would I check that on my own setup?

The sandbox holds. The model is what stops the leak.

· 8 min read
Pascal Nehlsen
Platform & Security Engineering

Last time I argued that everyone's comparing agent sandboxes on the wrong axis: the isolation boundary is largely solved, and the security question that matters starts after the wall holds. This is the first experiment from behind that wall.

A short field report from a lab I run on the security of autonomous AI coding agents inside microVM sandboxes. Everything below ran on my own hardware, against my own secrets and my own infrastructure. Half a dozen experiments in a day, and the interesting result isn't an exploit. It's where the control that actually stops a data leak turned out to live.

Everyone's comparing agent sandboxes on the wrong axis

· 2 min read
Pascal Nehlsen
Platform & Security Engineering

If you've looked into running AI agents safely, you've seen the comparison: Firecracker vs gVisor vs plain containers vs full VMs. microVMs boot in milliseconds, containers share the host kernel, gVisor sits in between; Firecracker powers a lot of it, and several products build sandboxes on top. The whole debate is about the isolation boundary: how hard is the wall between the agent and your host?

Here's what I keep running into: that wall is largely solved, and it's not where agents actually get dangerous.

Ephemeral AWS Sandboxes: 80+ Isolated Environments at Half the Cost

· 3 min read
Pascal Nehlsen
Platform & Security Engineering

Giving every learner a shared environment is cheap and miserable. One person breaks it and everyone is blocked. Giving everyone a fixed, always-on instance is clean and expensive. This post is about a third option: per-user, production-like sandboxes on AWS that spin up on demand, clean themselves up, and cost about half what fixed-size instances would.

SLO-Driven Automated Rollback: Let the Metrics Pull the Cord

· 3 min read
Pascal Nehlsen
Platform & Security Engineering

A deploy that breaks production at 02:00 shouldn't wait for a human to wake up, read a dashboard, and decide to roll back. If you can define what "broken" means in terms your monitoring can measure, you can let the pipeline pull the cord itself. This post walks through wiring observability into deployment so that an SLO breach triggers an automatic rollback.

Agentic DevOps Runbooks with a Human-Approval Layer

· 3 min read
Pascal Nehlsen
Platform & Security Engineering

"Let the AI fix it" is a great way to turn a small incident into a large one. But most of the toil in incident response isn't the fix. It's the gathering: pulling logs, checking deploy history, correlating metrics, reconstructing what changed. That part is safe to automate. This post describes a runbook executor that automates the gathering and the proposing, while keeping a human firmly in front of anything destructive.

Golden Paths on GCP: Cutting Provisioning Time by 80% with Terraform

· 3 min read
Pascal Nehlsen
Platform & Security Engineering

The fastest way to slow a team down is to make them wait on infrastructure. When every new service means a four-hour manual walk through the GCP console, engineers batch their requests, context-switch while they wait, and quietly build snowflakes. This post is about replacing that ritual with a golden path: a paved, opinionated route that provisions a production-like environment in under 45 minutes, self-service.