AI Fundamentals

Should You Build Your Own AI Data Analyst with Claude and MCP, or Buy a Platform?

Building your own AI data analyst with Claude and MCP is possible, and the connection itself takes an afternoon for a capable engineer. Licensing a platform means a vendor has already built, and keeps maintaining, the infrastructure around that connection: the semantic layer tooling, the permission-mapping framework, and the verification path that let your answers stay accurate as the business changes. Your definitions stay yours either way. What changes is who builds and maintains the machinery that holds them. The right choice depends less on whether your team can build it and more on whether anyone is named to own it after launch.

A data engineer on your team can connect Claude to your warehouse this afternoon. MCcP makes that part genuinely simple. Three months later, the same setup is quietly wrong more often than anyone's comfortable with, and nobody can say exactly when that started. This article walks through why that happens, what building it properly actually costs, and how to weigh that against buying a platform that has already solved it.

What Building an AI Data Analyst on Claude and MCP Actually Involves

MCP solves one problem: getting Claude to see your data. It doesn't solve whether Claude understands what that data means. That gap between what the schema says and what the business actually means is the semantic layer, and building one is ongoing work someone has to own.

OpenAI's own engineering team wrote in detail about building exactly this for their internal data agent (source). They built six separate context layers: table usage patterns and lineage, human-written annotations, code-level enrichment pulled from the pipelines that produce each table, ingested institutional knowledge, a persistent memory system, and live runtime queries for anything stale. Their own conclusion: "meaning lives in the code that produces it," not in the schema or query history. That was built by a company with more in-house AI expertise than almost any other, backed by a dedicated Data Productivity and Data Science team. That's the baseline the "have an engineer wire this up in an afternoon" plan is actually competing against.

Why the Demo Always Works and Production Doesn't

Anthropic published the numbers from building this exact kind of system internally (source). Without a semantic layer, Claude's accuracy on their internal analytics evals didn't exceed 21%. After building a proper semantic layer, that jumped to 95%+ in aggregate. Today, 95% of Anthropic's internal business analytics queries are automated through Claude at roughly that accuracy.

Then the part that matters more for this decision: when the team stopped actively maintaining that layer, accuracy drifted from around 95% down to 65% over a single month, before anyone treated it as an engineering problem worth fixing again. Nothing about the model changed in that month. The business did.

This is the same pattern that shows up across agentic workflows in analytics generally, not just at Anthropic. An agent that plans and executes multi-step analysis is only as reliable as the context feeding it, and that context has an expiration date unless someone is actively renewing it.

What the Maintenance Actually Costs

Anthropic names four ongoing components: data foundations, sources of truth, skills (procedural knowledge for how to approach a question), and validation. Each needs an owner, a standing responsibility, not a project plan with an end date.

The afternoon it takes to wire up MCP is real and small. The team-quarter it takes to build a semantic layer that doesn't rot, and the ongoing fraction of an engineer's time it takes to keep it from rotting, is the actual size of the "build" decision. That's the cost most teams underestimate when they scope this as a connectivity project instead of a standing one, the same underestimation that shows up in how self-service analytics projects get scoped more broadly: the connection is easy, the ongoing ownership is the actual project.

Build vs. Buy: A Direct Comparison

Both paths can get Claude or an equivalent model answering questions against your data. What differs is who does the ongoing work, and what that work costs over time.

Dimension Build (Claude + MCP, in-house) Buy (a platform built for this)
Time to first answer Hours to days, once MCP is wired to your warehouse Typically about a week to connect data and start running real questions
Who maintains the semantic layer infrastructure Your team builds and maintains both the tooling and the definitions, indefinitely The vendor maintains the infrastructure that holds the semantic layer; your business still owns the definitions inside it, and can edit them directly
Cost shape over time Low upfront, then an ongoing and often underestimated maintenance cost Predictable subscription cost, with infrastructure maintenance already priced in
Who maintains permission mapping Your team builds the mapping framework and re-maps it every time org structure changes The platform provides the infrastructure supporting the mapping; you keep your own roles and scopes current inside it
Verification / audit trail Something your team has to design, build, and maintain separately A built-in traceability path from any answer back to source, plus a human verification step, shipped as part of the platform
Business users defining their own terms Requires engineering involvement for every new term or metric A knowledge management layer where business users add terms and KPIs without a data engineer

Read row by row, the build column is entirely achievable. Every one of those is something a competent team can do. The difficulty is cumulative: all six, maintained at once, for as long as the system is in use. That is what makes the decision bigger than the afternoon it takes to get a first answer out of a working MCP connection.

When Building Genuinely Makes Sense

If you have engineers who can be assigned to this as an ongoing responsibility, not a side project, building makes sense. If your use case is narrow, the semantic layer you need is smaller and more tractable. If your end users are technically literate and can sanity-check a wrong answer, the cost of drift is lower.

Where it breaks down is the common case: business users who can't independently verify a query, questions that span the whole business, and no one whose job description includes keeping the semantic layer honest. That's also where a platform built specifically for operational business data, rather than a general-purpose model wired to a warehouse, tends to close the gap faster. Lumi AI is one concrete example of the buy path: a conversational, agentic analytics platform built for ERP, inventory, sales, and fulfillment data specifically, with a multi-agent architecture that plans and executes multi-step analysis, checks its own work, and explains its reasoning rather than returning a single query's worth of answer.

The knowledge management layer is the part that most directly answers the maintenance question above. Business users define their own terms, metrics, and KPIs inside the platform, without filing a ticket to a data engineer every time the business changes. Lumi's 3.0 release added persistent, editable, versioned chat artifacts and in-thread collaborative comments, so an answer doesn't disappear the moment the chat window closes, and a team can return to and build on a prior analysis instead of starting over.

On real deployments, Kroger used Lumi to de-average and re-aggregate data down to the store-item level, surfacing millions of units in unfulfilled demand. Chalhoub Group identified $60 million in additional revenue opportunity by driving in-store purchases. At GROWMARK, Jordan Kuhns described it this way: "If somebody has the concept of a KPI or metric that could help us make better decisions, Lumi can do all that work if they know to ask that question." Lumi also reports a chocolate manufacturer client released 15% of working capital by identifying unproductive and obsolete inventory using the platform; that figure is Lumi's own reported result, not an independently audited one, worth asking about directly rather than assuming as a baseline. Lumi also reports typical time-to-insight improving from seven days down to thirty seconds, and retail and CPG teams are typically live and running real questions within about a week of connecting their data. Lumi has completed a SOC 2 Type 1 audit, with Type 2 underway, and publishes live control monitoring through its Trust Center for teams that need to verify security posture as part of a build-versus-buy decision, alongside Lumi's broader security practices. Queries run against your data where it already sits, with no replication and no external storage of raw records.

None of this makes the build path wrong. It makes the buy path a real comparison rather than a shortcut, because the work a platform like Lumi has already done is the same work a build path has to do eventually: a maintained semantic layer, permissions that track org changes, and an audit trail for every answer.

A Real Checklist Before You Decide

  1. Who is the named owner of the semantic layer once this is live, and what percentage of their time is actually allocated to maintaining it?

  2. What's your plan for detecting accuracy drift before a business user does?

  3. Can every answer be traced back to the exact query and rows that produced it?

  4. How will permissions be mapped and kept current as your org structure changes?

  5. If you're buying instead, how was the vendor's accuracy measured, on what kind of questions, and how recently? Lumi's own accuracy evaluation is a useful reference point for the kind of answer a vendor should be able to give you.

  6. Does your data team want to spend its time on the analysis itself, or on building and maintaining a custom application layer on top of an API? That question alone often settles the decision for teams already stretched thin, which is also why data teams evaluating this tend to weigh their own capacity as heavily as the technology itself.

If your own team or a vendor can't answer all six specifically, you're not evaluating a finished system. You're evaluating a demo.

FAQ

Is MCP hard to set up for connecting Claude to a data warehouse?

No. A capable engineer can typically get Claude answering real questions against a live warehouse within a day. The difficulty is everything that determines whether those answers stay accurate afterward.

What's the actual difference between build and buy here, if both use similar AI models?

The models are the commodity. The difference is who builds and maintains the scaffolding around them: the semantic layer, the permission mapping, and the verification path. Building means your team owns that scaffolding indefinitely. Licensing a platform means a vendor maintains it for you. To be precise about what that does and doesn't cover: Lumi does not maintain your business context for you. Your metric definitions, exclusions, and terminology remain yours. What Lumi maintains is the infrastructure that makes keeping them current a task a business user can do directly, rather than an engineering ticket, plus the evaluation framework that tells you when a definition has drifted.

How do I know if our internal AI analytics setup has started drifting?

Continuous validation, not a one-time launch check. Anthropic's own reported drift from roughly 95% to 65% over one month, unnoticed until it became a visible problem, is a reasonable model for how invisible this can be without dedicated monitoring.

Does a small pilot project justify building this in-house even if a full rollout wouldn't?

Often yes. A narrow, well-defined use case with a technically literate user base carries much less semantic-layer risk than open-ended, business-wide analytics.

What does a platform like Lumi AI actually take off my team's plate compared to building it ourselves?

Building and staffing the machinery: the semantic layer tooling, the permission-mapping framework, and the verification path a build path has to create from scratch. It does not take your business context off your plate, and no platform can. What Lumi's conversational analytics interface and knowledge management layer change is who can keep that context current: a category manager or planner updating a definition directly, rather than an engineer working through a ticket.

How long does buying typically take before a team is running real questions?

Lumi reports that retail and CPG teams are typically live and running real questions within about a week of connecting their data, compared to the team-quarter or longer it can take to build and validate a semantic layer in-house.

Where This Leaves You

MCP genuinely solved the connection problem. What OpenAI's own engineering writeup and Anthropic's own accuracy numbers both confirm is that the part after the connection is a standing job, not a project. Buying a platform doesn't remove that job, it means someone else is already doing it, and has been for longer than your team has had this on their backlog.

The readiness bar that job actually has to clear is the same one either path has to meet: the same question should return the same answer regardless of who asks it or how they phrase it, checked continuously, not once at launch. If you want to see what that looks like against your own data before committing engineering time either way, schedule a demo or review Lumi's pricing to compare the cost of buying against the real, ongoing cost of building it yourself.

Social Media
Ibrahim Ashqar

Data & AI Products | Founder & CEO at Lumi AI | Ex-Director at Unicorn. Ibrahim Ashqar is the Founder and CEO of Lumi AI, a company at the forefront of revolutionizing business intelligence for organizations with a specialization in the supply chain industry. With a deep-rooted passion for democratizing data access, Lumi AI seeks to transform plain language queries into actionable business insights, eliminating the barriers posed by SQL and Python skills.

Lumi AI Connection Graphic for Analytics 101 blog page sidebar

Illuminate Your Path to Discovery with Lumi

Explore Pilot Program

Related articles

The New Standard for Analytics is Agentic

Make Better, Faster Decisions.

Request Demo
2026-09-16
2026-09-16