Context Layer vs File Search: A Private Equity AI Benchmark

Taylor Lowe

Key takeaways

We asked the same AI model 51 real-world private equity questions twice, across the same 100,000 files: once connected to Metal, once to a leading enterprise file platform, each through its own MCP server.

Same questions, same data, same model. The only difference was the path to the data: through Metal's context layer, or straight to the raw files.

Metric

Result

Detail

Quality

95.1 vs. 86.6

Metal scored 95.1 out of 100; the file platform scored 86.6.

Hardest questions

Up to 50 points higher

The gap was widest on complex work, such as reconciling conflicting EBITDA versions, building financing chronologies, and tracing covenants.

Reliability

3.4× fewer errors

The file platform produced 3.4× as many tool errors per question (9.1 vs. 2.7), and 13.5× as many on structured workflows (18.9 vs. 1.4).

Cost

41% fewer tokens

Metal used 1.13M tokens per answer, vs. 1.90M on the file platform.

Speed

2× faster

Metal answered in 3.8 minutes on average, vs. 7.6 on the file platform.

Background

"We already use Claude or ChatGPT. Why do we need a context layer?"

It is one of the questions we hear most from private equity firms, and it rests on a reasonable assumption: answer quality is mostly a function of the model. If the model is strong enough, pointing it at the firm's files should be enough.

Private equity work rarely turns on finding a document. It turns on conviction, and conviction is built from several things at once: the firm's history with a company, the framework it underwrites against, the version of each number that governs, the evidence behind it, and the argument that ties them into a decision. A file search can return the document. It cannot bring the rest together.

Existing finance benchmarks, such as FinanceBench and FinQA, say little about what happens when a firm connects a model to its own private record, where versions conflict, dates are ambiguous, and state lives outside the documents.

So we ran the opposite experiment: hold the model constant and change only the data layer it reads through. 

How We Ran the Evaluation

Test Setup (Learn more)
  • Model: one general-purpose agent running GPT 5.6 Sol, with identical instructions on both sides.

  • Backends: the agent reads through one of two backends, each connected over MCP. One was Metal's context layer. The other was a leading enterprise file platform's own MCP server, which is how most firms connect AI to their documents today.

  • Dataset: a real-world private equity dataset. 

  • Output: alongside quality, we recorded the work behind every answer, covering time to answer, total tokens across agent turns, and failed tool calls.

Quality Score Across Three Question Banks

Each answer was scored from 0 to 100 by model judgment with source review. Structured workflow scores are the average of factual accuracy and task completion.

Metal scored higher in all three banks. The margin was widest on the Hard bank, 96.1 against 83.1, where questions depend on reconciling versions and periods. It was narrowest on CIM discovery, 95.4 against 91.7, where most answers sit inside a single set of sell-side materials.

Where Does the Gap Open Up?

Metal scored higher on 35 of the 51 questions, tied on 5 and scored lower on 11. Nine questions with margins of 20 points or more account for about 70% of the total gap. Across the other 42, the average margin is 3.2 points.

The nine questions share a pattern. Each answer depends on something no single document states: version, period, entity, or deal state.

Scenarios

What breaks for data file platform

What's needed

Version

A forecast, several actuals, a draft and a final all match the same search. Nothing in a file index says which one controls, and an answer built on the wrong version is wrong in every figure that follows.

Version lineage: which document supersedes which, and which version the firm reports from.

Time

A number is only correct when it is tied to its period and definition. Sourcing has the same problem: what matters is when the team received a CIM, not when the file was created or imported.

A period and definition attached to every figure, and process dates stored separately from file dates.

Identity

A software vendor serving hospitals matches every healthcare search without being a healthcare company. And one business spread across a CIM, a teaser and several copies gets counted several times.

One record per business, with its own classification, linked to every file that mentions it.

State

Whether a screening exists, is linked to the right company, or finished processing is recorded in the pipeline, not in any document.

The pipeline itself: records, statuses, and the links between company, screening and source.

The gap opens where the answer depends on firm context rather than on a single file. The agent performs better with Metal context layer when it requires knowing which version governs, which date is real, or what state a deal is in. The file platform has the documents but not the structure that connects them.

Where the File Platform Led

The file platform scored higher on 11 questions. Ten of those margins were 6 points or less. The exception was a comparison of customer concentration and data-center exposure across three power-services companies, where the file platform scored 99 and Metal 85.

Most of these questions name the companies up front and ask for a comparison whose evidence sits in those companies' own materials. Examples are trade-data defensibility (100 against 94), retention-metric comparability (88 against 84) and a quarterly reconciliation for a single platform company (92 against 89). When the businesses are named and the answer lives inside their documents, a search-and-read loop over good files gets there. That is the expected result, and it is why Metal sits on top of a firm's file platform rather than replacing it. The context layer earns its place on questions whose answer no single document states.

Tool Errors: Where the Agent Stalls

Across all 51 questions, the agent averaged 2.7 tool errors per answer through Metal and 9.1 through the file platform, 70% fewer. The difference is concentrated in one bank:

Our reading is that the 18.9 sits where the State gap sits. Workflow questions need pipeline state, and an agent that looks for statuses and links that a file store does not hold keeps making calls that fail. The Hard result shows that an error is not a wrong answer, because an agent can recover from a failed call. For teams running agents on their own stack, the error rate is still a reliability signal. A workflow that errors 19 times per answer is hard to run unattended, and every failed call still costs a turn and tokens.

Token Cost per Answer

Better models take on bigger jobs, and bigger jobs burn more tokens. Teams running agents on Claude and ChatGPT keep telling us their token bills climb every month. So we compared what each backend cost per answer.

What This Means for Firms Building Their Own AI

Most firms building in-house start with workflows: an agent, a set of prompts, and a connection to the file system. It feels like progress, because a good model pointed at files produces a convincing first answer. This test shows why that order is backwards. The model, the agent and the questions stayed the same, and the results changed only when the data layer underneath changed. A workflow is only as reliable as what it reads from.

  • A workflow handles one kind of task. It decides what to ask, which tools to call and what to return. It can be scoped, built and shipped as a project.

  • A context layer is what every workflow reads from. It resolves each company across aliases, codenames and duplicate records. It knows which version controls and when a document was received, not only when it was created. It holds pipeline state alongside the documents. And it has to stay current as deals close, documents are revised and people change seats. That makes it infrastructure the firm operates, not a project it finishes.

Without that layer, each workflow rebuilds the same structure on every query. That is where the nine largest gaps in this test came from, and it is what the token and error counts measure. It also means two workflows can resolve the same company, or the same EBITDA, in different ways. The deal team, the portfolio team and IR then end up with answers that don't match. This is where context graphs move beyond RAG.

For private capital firms serious about AI, the most important decision comes before the first workflow: what they will all read from. When every agent works from one record, each new workflow starts from a structure that is already resolved, and the whole firm can rely on the answers.

What Are We Testing Next?

This is the first result we are publishing. We are running the same method across more question banks, more models, and more of the workflows firms run every week, and we will share those results as they come in.

The most useful version of this test is the one run on your own data. Talk to the Metal team.


Join top firms redefining private capital with AI

Join top firms redefining
private capital with AI

Join top firms redefining private capital with AI