Project · Design and build
OpenFlow Metrics
A flow metrics and forecasting add-on for Jira, given away rather than sold. Thirteen charts, a promise a team writes down and is then held to, and forecasting that answers "when" with a date and a confidence instead of a number of points. None of the arithmetic leaves the browser, and that one decision is why it can be free.
Reading level
The plain-English version. Every section also has a technical note.
What it is
An add-on that reads a team's own tracker and draws how work actually flows through it — how long things take, how much gets finished, and what is stuck right now.
There are thirteen charts. Most of them describe work that has already finished: how long items took and how variable that was, how much the team completes each week, where work piles up, how much of an item's elapsed time was spent working rather than waiting, and whether a change in any of those is a real signal or ordinary noise.
The view it opens on is none of those. It opens on ageing work in progress — the items in flight, ranked by how long they have been going. That is the only chart in the set that describes work anybody can still do something about, and putting it first is the whole editorial position of the tool. A team that reads it in a stand-up has a list of things to act on; the rest is for a retrospective.
The two forecasts, and the trap between them
Forecasting is done by sampling the team's own completed history thousands of times over rather than by adding up estimates. It answers two questions: when will these items be done, and how many will be done by a date. Confidence runs in opposite directions across the two, which is the thing most likely to be read backwards. For a date, higher confidence means a later answer. For an amount, higher confidence means a smaller one — to be sure of delivering at least a number, that number has to sit low in the distribution. Getting it the wrong way round turns a forecast into a trap, so the two are worded and tested as opposites rather than as one shared control.
Why it is free, and what that decided
The platform bills the person who wrote the add-on rather than the team using it, so an add-on that did its thinking on a server would cost me more the more people liked it.
That is the constraint the whole design comes out of. The obvious build — fetch every item's history on a server, compute the metrics there, send back a chart — puts every customer's usage on my bill with no revenue against it. Free would last exactly as long as nobody used it.
So it does not do that. The page in the browser talks to the tracker's own interface directly, and every calculation happens on the reader's machine. The server side is deliberately tiny: it remembers which columns count as started and finished for a project, and it answers whether the licence is valid. Nothing else.
Three things fall out of it
Running cost stays near zero however many teams install it, which is what makes free sustainable rather than a promise with a shelf life. There is no server-side time limit to design around, because nothing heavy runs on a server — a team with years of history is a slower page, not a failed request. And the issue data never leaves the vendor's own walls, which is both the honest privacy answer and the qualification for the platform badge that says so.
What it cost
Everything has to be re-fetched and recomputed in a page that starts from nothing, so the work went into asking for less rather than into caching more. The arithmetic itself is a layer that knows nothing about the tracker, the platform or the interface framework — plain functions over plain data, with no dependencies of their own. That is what makes it testable, and the arithmetic is the part that has to be right.
The promise, and the list it produces
A team writes down what it expects of itself — "85% of our work finishes in eight days or less" — and the tool holds it to that, then tells it which work is about to break it.
Every tool of this kind can draw a percentile line across a chart. What none of the free ones does is treat that line as a thing the team owns: saved, revisited, measured against, and turned into a list. Saying the sentence out loud is what makes it actionable. A line on a chart is a fact about the past; a written-down expectation is a promise, and a promise can be checked.
What comes out of it is a ranked list of the work in flight that is going to miss. Items already past the expectation are marked as breached; items past roughly two-thirds of it are flagged while acting is still cheap. Worst first, with who has it and how far over it is. That list, not the charts, is what a stand-up can actually work through.
Two decisions that keep it honest
Where a team has not set one, an expectation is derived from its own finished work and marked as derived rather than presented as a choice the team made. Deriving is refused outright below ten completed items: an authoritative-looking promise computed from four data points is worse than no promise at all.
And when a team is routinely missing its expectation, the tool says so in those terms — an expectation that is missed most weeks is a measurement that was set wrong, not a team that is trying insufficiently hard. A metrics tool that lets itself be read as a stick is a metrics tool nobody will keep open.
Why the percentile is a rank and not an average
Percentiles are taken by nearest rank rather than by interpolating between two items. Nearest rank returns a duration something actually achieved, so "85% finish in eight days or less" is literally true of the data. Interpolating would return a figure like 7.4 days, which no item ever took, and the promise would quietly stop holding while still looking precise. Several conventions are pinned the same way and each has a test on it: a day is counted whole and inclusively, so work started and finished in one day is one day rather than none; an item bounced back to an earlier column does not get its clock restarted; and an item reopened after finishing counts as in flight again rather than staying done.
The demo, and what is in it
There is a public demo you can open now, and every number in it is invented.
It runs the real add-on — the same charts, the same arithmetic — against a tracker that is generated in your browser when the page loads. No account, no install, and no real team's data anywhere near it. Ten scenarios can be switched between, because a healthy team makes a boring chart: the one worth opening first is the struggling team, where the ageing list has something in it and the expectation is being missed.
Why synthetic data rather than a sample export
A fixed export would be one team's shape for ever, and the charts that matter here are the ones that change between a team that is fine and a team that is not. Generating the history in the browser means the scenarios can be switched while the page is open, and it means the demo carries no consent question at all — there is nothing in it that belonged to anybody.
How it is checked
The arithmetic has more than three hundred tests on it, and the build fails if the writing that explains a feature falls behind the code.
The maths layer is where the tests are concentrated, because it is the part that can be wrong without looking wrong. A percentile off by one rank still draws a plausible chart. So the conventions above are each pinned by a test that fails if somebody decides the other reading was more natural.
The part I would defend hardest is not the coverage number, though there is a committed floor and the build refuses a drop below it. It is that every feature has a written specification and at least one decision record, tied to the source files that implement it, and a check that fails the build in two directions: a source file belonging to no feature, and a feature whose code changed while its writing did not. Twenty-two features are registered that way, with twenty-seven decision records behind them.
Why documentation is a gate rather than a good intention
Every project I have worked on has had stale design documents, and the reason is always the same: nothing fails when they go stale. Making it fail is the only mechanism I have found that works. The cost is real — a change to a feature is a bigger commit than it would otherwise be, and there are days when that is annoying. The benefit is that the reason a percentile is computed one way and not the other is still written down months later, which is exactly when somebody is going to want to change it.
Where it has got to
It is built and it runs, but it has never been installed on a real team's tracker and it is not on the marketplace.
Everything above is true of the code and provable from the demo. What is not yet true is any of the things that only contact with a real installation can settle: whether the column-mapping guesswork survives a workflow somebody grew over five years, how it behaves against a project with tens of thousands of items in it, and whether the reading it offers changes what a team does. Listing it also has requirements of its own that are not met yet.
So the honest description is a finished argument with an unfinished proof. The next step is a real installation on a real workflow, and until that has happened the demo is the whole of the evidence.
What I expect to break first
Not the arithmetic — that is the best-tested part and the easiest to reason about. It will be the workflow detection: deciding which of a team's columns mean started and which mean finished, on a board whose statuses have accumulated for years and whose names agree with nothing. Every metric in the tool hangs off that mapping, and a mapping that is subtly wrong produces charts that are perfectly plausible and quietly false.