How we built the IT Process Graph
Blog/Technology

How we built the IT Process Graph

The IT process graph is a map of how the business runs across its systems, built from the work itself and kept current as the systems change. This is how building our agentic harness taught us we needed one.

In The CIO’s third act, Rahul Kayala explains why enterprise AI programs rarely stall on model capabilities. They stall because the AI has no real understanding of the systems underneath it. The knowledge of how work moves across a company's systems used to sit with IT. Over twenty years, it scattered across SaaS apps, partners, and people's heads.

At Echelon, we’re building two things that work together. The first is an agentic harness for enterprise applications, so what an IT team can take on is no longer capped by the hours its people have. The second is the IT process graph, which gives the harness the context it needs to do the work well. It's a map of how the business runs across its systems, built from the work itself and kept current as the systems change.

This post is about the IT process graph, and how building the harness taught us we needed one.

Every task started cold.

Picture an IT team that owns the systems below and the processes that run through them. Nothing is documented. No one person knows a whole process, or even every system it touches.

So every bug or troubleshooting request on a system of record becomes a ticket, and every ticket starts with an investigation: query the systems one by one until the cause turns up. Then the fix. Then the next ticket, and the same investigation again.

A company's systems fill the frame: Workday, Fieldglass, Okta, ServiceNow, Jamf, Microsoft 365, Slack, Salesforce, SAP S/4HANA, Coupa, Stripe, DocuSign, Zendesk, Jira, Snowflake and BlackLine. Its end-to-end processes run across them, and each one lights up the systems it touches: Order to Cash, Source to Pay, Case to Resolution, Request to Fulfillment, Record to Report, and Hire to Retire, which stays lit while the rest dim. Security files a ticket: contractors keep access after they leave. The agent investigates system by system: in ServiceNow the user record is still active; Okta has no deactivation event; the person was never in Workday; Fieldglass has the end date but no feed to Okta; back in ServiceNow, the group-removal flow is still a draft. Resolved. Root cause: Fieldglass has the end date, and nothing sends it to Okta. Then a finance owner reports a bug: someone who left still approves their purchase orders. The agent starts in ServiceNow, where the person is still in 4 groups, one of them approvers, then walks the same way again: the draft group-removal flow, the user record, Okta, Workday, 4 of its 5 steps a repeat. Resolved. Root cause: group removal is still a draft, so they stay an approver. What both tickets walked lines up as one row, the Leaver stage of Hire to Retire: end date recorded (Workday or Fieldglass), account switched off (Okta), groups removed (ServiceNow), each marked by both. A ticket and a bug report needed the same process. Nobody had written it down, so each one pieced it together from scratch.

When we first built our harness, it took on that investigation. It can analyze, design, configure, troubleshoot, and document these systems. It found the cause and applied the fix. But the next ticket started from nothing, just like the last one. If it had kept what it learned from that ticket, the next one would have started at the group-removal flow instead of at the first system. What it learned is the knowledge Rahul says never compounds.

Everything is a data point to learn from

The fix seemed obvious: let every task write what it learned to a shared memory that outlives the task.

A page of memory, labeled. The page is called Access removal. At the top, its state says how settled the page is: first piece, partly known or complete; this one is partly known. The body is four facts, each one thing a task showed about this company, and each with a source line: the task and transcript line it came from, and how it is known, seen in a system or said by a person, set by code from the source line, not by the model. Then two numbered open questions, each naming who could answer. Answered questions are struck through, never deleted.

Memory is a collection of pages, like the one shown above. They capture the “decision traces” that used to live in people’s heads.

  • Each “page” is one short document, about one thing: a system, a process, a team’s convention, or a piece of work in flight.
  • A page is made up of “facts”, something that a task showed about the memory, such as “a contractor’s access ends on their end date”.
  • Facts are attributable to their “origin” and have a label on “how the fact was determined”. So a fact said by a person is much stronger than something found on the web, or things the agent concluded.
  • “Open Questions” are something that a page doesn’t know yet - something that a future task could answer.
  • Finally, “State” captures how settled the content of the page is - transitioning from first touch to complete.

We could have started with a more structured organization: a page for every system, a page for every process, fields for owners and approvals, etc. But we chose not to. Every organization works differently, and a fixed schema chosen up front would describe how we think companies run, not how this one does. So we designed only what a page holds, and let the work decide which pages exist.

The memory we curate comes from two places.

The first is the everyday work people run through Echelon: troubleshooting a bug, answering a question, or changing the systems themselves. Each task has to understand the process before it touches it, and when it finishes, it writes down what it learned.

The second is configuration present in the system. Configurations reflect the reality of the business process, not what the old design document claims. History shows who changed it and why: a rule added by a contractor who has since left, custom code rushed in for a go-live, the same fix rebuilt by one integrator after another because none of them knew the last had been there.

What learning looks like

A task writes to memory as it works. Here’s the finance bug again, showcasing how the memory gets updated based on what the harness found.

A finance owner reports a bug: someone who left still approves our purchase orders. The agent works it across the company's systems, and writes to memory as it goes, into the file access-removal.md, which already holds a line from Security's ticket: group removal in ServiceNow is still a draft. In ServiceNow, group removal is still a draft; memory already has this, so the bug report is added as a second source on that line. Still in ServiceNow, the person is in 4 groups, one of them approvers; new, so purchase-order approvers are a ServiceNow group is added. In Okta, the deactivation came from Workday; new, so Okta switches accounts off when Workday records a termination is added. In Workday, the termination is recorded for this person; it is only about one person, so it is not kept. Resolved. What the bug report wrote sits in one yellow cell in the memory file, each line with the task and line it came from.

Three things can happen to what a task finds. If memory already has it, the task adds itself as another source, and the fact gains credibility. If it's new and would change what a future task does, it's added. If it only matters today, like one person's termination date, it's dropped. The source line on every fact is recorded by code, straight from the transcript, never written by the model.

The shape that memory took

We modeled the shelves after what a forward-deployed engineer’s notebook would look like after a few months in an organization.

The product already kept memory in a few shapes (vocabulary, how things are set up, workflows, work in flight), so those became the shelves. Beyond that, we designed nothing. Which pages existed, and how they grouped, was left to the harness.

Pages of memory piling up over three weeks, one square per page, each landing in a column for its kind. Nobody set the kinds up front: each kind's name appears only when its first page does. How things are set up, vocabulary and work in flight appear in the first days, step-by-step workflows right after. Recall notes appear only in the second week, once there was enough memory to need a way into it, and end as the largest kind. After three weeks, about 140 tasks have made about 60 pages: how things are set up 15, step-by-step workflows 11, vocabulary 9, work in flight 9, recall notes 16. Numbers are illustrative and rounded.

A lot has been written about agents creating memory for themselves. We think recall is equally important - if not more so. Memory a task can't find is memory it doesn't have. So the harness writes recall notes, each describing how a kind of task tends to open, what to settle first, and which few pages to read. A new task starts with the slice it needs, not the whole memory.

Just memory isn’t enough

Memory worked. Tasks found what earlier tasks had learned, and repeated questions got shorter answers. But it continued to grow.

One page of memory, Access removal, over three weeks of work, with every line drawn as a bar. It starts with 4 lines and 3 open questions. Every day new tasks add lines and questions, and the page gets longer. In the second week a correction lands: removal runs at 03:00 since the job moved. It sits beside line 3, removal runs nightly at 02:00, and nothing replaces the old line, so both are kept; both turn red. Beside it, three more memory files, workday.md, contractors.md and offboarding.md, fill in at the same time. In the third week the same fact shows up in four versions on this page and one more in each of the other three files, seven in all, each updated on its own: line 1, Workday is the source of truth for end dates, climbs from agent concluded to person agreed to said by a person; line 7 gains two sources; line 18 one; line 34 is new. None has the whole picture; all seven turn red. Then one transcript line turns out to be the quote behind about half the lines on the page, and on the other files too. A weekly clean-up rewrites the page three times, and each time it comes out one line longer. After three weeks: 37 lines, 8 open questions, and the page is still partly known, never complete.

When the removal job moved from 02:00 to 03:00, the new line landed beside the old one, and nothing retired it. The same fact about Workday lived on several pages, each copy updated on its own. Open questions piled up, and no page ever reached "complete."

We also found the agents gaming our own check. Every fact had to cite the line in the task it came from, so the model found a line that passed and cited it again and again. On one page, most facts pointed back to the same line, which couldn't explain them all. Every citation checked out, and most of them meant nothing. We stopped treating a passing citation as proof, and started judging pages by whether a new engineer would use them.

The deeper problem was the container. In text, a source, a correction, and an open question are all just words. Nothing marks two lines as the same fact, or says that one replaces another. Tidying up meant a model rewriting whole pages, and the cost of that grew with every task.

Let there be process graphs!

What if a fact weren't a sentence on a page, but a record? There would be a record for Workday, one for Okta, one for the group-removal flow.

Each fact about them keeps its sources and how it's known, exactly as before. The difference is that the structure the text was hiding is now something the harness can work with.

Three things change.

  • A fact is one record. When four tasks say where end dates come from, the record gains four sources, instead of the page gaining four lines.
  • New values replace old ones. 03:00 replaces 02:00. The old value stays in the record's history, with the task that said it, but nothing reads it as current.
  • Records link. Workday feeds Okta, Okta feeds ServiceNow, and ServiceNow runs group removal. Following the process, those links are the IT process graph for the “Leaver” stage of the Hire to Retire workflow.

The links also show the gaps: a missing link or a flow still in draft.

How the graph comes to be. It starts from the Access removal page as figure 5 left it: 37 lines of text, some of them red because they repeat or disagree. The lines lift off the page and gather into records, one for each thing they are about: Workday, Fieldglass, Okta, ServiceNow and group removal. The four lines that said where end dates come from become one record with four sources. The 02:00 and 03:00 lines become one: 03:00 replaces 02:00, and the old one stays in history. Then the records link: Workday feeds Okta, Okta feeds ServiceNow, ServiceNow runs group removal. The gaps show up on their own: nothing links Fieldglass to Okta, and group removal is still a draft. Then the view pulls back: the linked records fold into one stage, Leaver, of Hire to Retire, still marked with its 2 gaps, in a map of the company's processes, Hire to Retire, Order to Cash, Source to Pay, Case to Resolution, Request to Fulfillment and Record to Report, each with four stages. 8 of the 24 stages are captured; then more land one at a time, each with a pulse, until 14 of 24 are captured and counting.

Facts still originate from tasks, and every record still says where it came from. What changes is who keeps them tidy. Finding a fact across tasks, retiring a stale value, and closing a question become deterministic graph updates, instead of a model rewriting pages.

Every task starts warm

We started with a promise: what an IT team can take on should no longer be capped by the hours its people have. The harness alone got us partway there. It could do the work, but every task paid for its own discovery.

With the process graph, a task starts where the last one stopped. That makes it faster, but more importantly, it shows what comes after the fix. IT used to know that. The graph brings it back: the systems that feed each step, the ones that depend on it, and the people whose work changes with it.

The same bug report, run twice. A finance owner reports: someone who left still approves our purchase orders. Without the graph, the agent walks it the way it did in the opening: ServiceNow three times, then Okta, then Workday. 5 steps, about 35 minutes. With the graph, the agent first reads what is already known about the Leaver stage of Hire to Retire: Workday switches accounts off in Okta, Okta deactivates the user in ServiceNow, ServiceNow owns group removal, and two gaps are marked: Fieldglass has no link to Okta, and group removal is still a draft. The draft is the gap this report is about, so the agent goes straight there: in ServiceNow the person is still in 4 groups, one of them approvers, and group removal is still a draft, as the graph says. 2 steps, about 5 minutes. Root cause: group removal is still a draft, so they stay an approver. The process was already in the graph, so the bug report went straight to the gap.

Every task leaves the graph more complete, so the next one starts further ahead. An IT team no longer has to choose which few changes it can afford to carry end to end.

The process, as it runs today.

The graph the harness works from is the same one the people responsible for a process need. Documentation shows how someone designed the process at some point. The graph shows how it runs now.

One graph, two views. First, the full view: the company's six business processes, each a card with the systems it runs across, the stages captured so far and its open gaps. Hire to Retire, 2 gaps. Order to Cash, 1. Source to Pay, 1. Case to Resolution, none. Request to Fulfillment, 1. Record to Report, none. Drilling into Hire to Retire shows its stages: Joiner, not yet captured; Onboarding, in Okta and Jamf; Mover, in ServiceNow; and Leaver, which opens into its steps: end date recorded in Workday or Fieldglass, account switched off in Okta, groups removed in ServiceNow. Two gaps are marked: contractors' end dates have no link to Okta, and group removal is still a draft. Security's ticket and Finance's bug report found them. Second, a planned change, described by an IT director: we're moving from Okta to Microsoft Entra ID, what changes? The graph is cut to that change, across processes, showing every step that runs on Okta. In Hire to Retire onboarding: hire recorded in Workday, account created in Okta, laptop shipped in Jamf. In Hire to Retire leaver: end date recorded in Workday or Fieldglass, account switched off in Okta, groups removed in ServiceNow. In Request to Fulfillment: access requested and approved in ServiceNow, access granted in Okta. The three Okta steps move to Entra ID: 3 steps across 2 processes. The known gap moves with them: contractors' end dates in Fieldglass have no link to Okta, and won't reach Entra ID either unless it is fixed.

Gaps in the graph are where the process actually breaks, found by real work. They are a natural place to start a redesign.

Every transformation starts with understanding how things run today. That step is slow, expensive, and quickly outdated. The process graph makes this easy by staying current - as everyday work keeps it updated.

It also removes the reason to think small. Most transformation plans cover the top few systems, not because the rest don't matter but because nobody can afford to understand them all. With the graph, the whole estate is in view, so a plan can cover every application a process touches, not just the ones someone had time to map.

The graph's projection isn’t fixed either. A stakeholder could bring a planned change to the graph to see every step it would touch, across every process, and which known gaps would move with it.

Book a demo to see the full cycle in practice.

Written By

Anand Sainath

Get Started

See Echelon build in your instance.