Here is the situation, in general. You’re mid-task in an AI coding assistant. You send a request. Nothing happens for eight seconds, then text pours out, then it freezes for four, then finishes. Was that your Wi-Fi? Your VPN? The provider having a bad afternoon? A tool the assistant was running on your own machine? Or just a long answer?
Every one of those feels identical from the chair. And none of the usual tools can tell them apart.
Why the usual answers don’t work
Ping and round-trip time tell you about the network, and only the network. A connection can have a perfect 20 ms RTT while the provider takes eight seconds to start answering. RTT alone can’t separate “my network is slow” from “their service is slow”, and that is exactly the question.
Browser dev tools are great, but only for AI in a browser. A growing share of AI use isn’t: Claude Code, Codex, the desktop chat apps, Copilot inside an editor. Native apps have no dev tools panel, and they keep their own timing to themselves.
Decrypting the traffic (a local proxy with your own certificate authority) would show everything, but it means installing a root certificate and reading prompts and answers. For a tool you want to run all day, on a work machine, with admin rights, that is the wrong trade.
The bet: timing is enough
The idea behind watson is that you don’t need the content to know how a response behaved. Encrypted traffic still has a shape: when your request left, how long until the first byte came back, how steadily the answer streamed, where it paused, whether packets were being resent while it did.
So watson watches packet timing and sizes, and nothing else. It reads one thing from the application layer, the server name in the TLS handshake, to know which service a connection talks to. It never decrypts, never stores content, and nothing leaves the machine. That promise is the reason it is acceptable to run with the privileges packet capture needs, so I wrote it down as a design rule and treat anything that would break it as out of scope.
From there it does what a person staring at a packet trace would do, automatically:
- finds each interaction, one request and its streamed response, from timing alone;
- measures time to first byte, how long the response took, and how much of the delay was connection setup;
- explains each pause longer than a second. If packets were being retransmitted or round-trip time spiked, it was the network. If your CPU was saturated, it was your machine. If nothing else was going on, it was the provider working, and that one is reported but never held against the network.
That last rule matters to me. A tool that blames the network every time an answer is slow is worse than no tool, because it sends you off to fix the wrong thing.
Being honest about what it can’t know
Inference has limits, and I tried to make the tool say so instead of hiding them. Every number carries how it was obtained: read directly from TCP state, inferred from timing, or not observable at all. QUIC traffic hides the connection’s health, so a pause there is labelled “unattributed” rather than guessed at. A score comes with a confidence rating and the reasons it is low. And what each interaction is judged against is your own history with that service, not a made-up threshold.
It also can’t see tokens, model names, or what you asked. I think that’s the right place to stop.
What it looks like now
You run watson watch in one terminal and it prints a row per AI response: time to first byte, duration, pauses, and a 0 to 100 transport score. Interactions are grouped into turns, so you can see how long a single “waiting on the AI” episode took and how much of it was actually on the wire.

In another terminal, watson serve opens a live page: a tile per metric against what’s normal for you, and a chart of every interaction’s time broken into setup, round-trip, server time, provider wait, network stall and transfer, with the score drawn over it. You can pick one app or all of them, over the last 12, 24 or 48 hours.

What I got wrong, and what I’d do differently
I blamed the network first. When a session feels slow, the Wi-Fi is the obvious suspect, and that was my assumption too. The data disagreed. Across two days and about 960 interactions, the median round trip was 23 ms, and only 7 interactions had a network stall. The network was rarely the answer.
A lot of the wait wasn’t mine to fix at all. Provider wait, where nothing was wrong on my side and the service was simply working, made up about 18% of the time spent inside responses. It’s reported but never counted against the network, so it can’t send me chasing the wrong problem.
And some of the wait wasn’t the network or the provider. In the turn pictured above, 32 seconds from start to finish, 24.7 of them were on the wire and 7.6 were “other”: tools running on my machine, the model thinking before it sent, and me.
That changed how I react when a session feels slow. I now ask the same questions in the same order:
- Is the network unhealthy? High round-trip time, packets being resent, or a pause that lines up with a retransmit means it’s mine to fix: switch networks, check the VPN, move closer to the router. This is the only case where the fix is a network fix.
- Is the network fine but the service slow? A long time to first byte or a long provider wait on a healthy connection means the delay is on their side. There is nothing to fix locally. I retry later, ask for something smaller, or try another model. Most importantly, I stop restarting my Wi-Fi.
- Are both fine and it still feels slow? Then the time is going to the part watson calls “other time”, which is how I’m using the tool. A turn full of tool calls, or a very long conversation, costs time that no network can explain. This is where I’d suspect my context window or the shape of the task, and the response is to break the work into smaller steps or start a fresh session.
I got a few things wrong while building it, too, and each one taught me something about the same lesson. My first chart of where time goes stacked the median of several unrelated measurements, which added up to something that looked authoritative and meant nothing, so I rebuilt it from parts that really sum to the interaction’s duration. A “throughput” tile turned out to mostly measure how fast the model types, not how fast the network is, so I replaced it with round-trip health. And a chart I had declared finished was completely blank in a real browser because of one unquoted attribute; it passed every check that read the page’s text. I now render and look.
Why bother generalizing
Streaming AI is the case that made the problem obvious, but the same question comes up for every desktop app: file sync that feels stuck, a softphone that cuts out, a remote desktop that lags. Each needs different metrics, because users feel different things. So the next step is a small set of profiles (request/response, bulk transfer, real-time media, remote session, push, AI streaming) that decide which measurements apply and how they are scored, never blended into one meaningless number.