◎entityfactsguide942.solsticebrief.com

Using MCP for Wikidata for Evidence-First Record Resolution

Record resolution sounds tidy on paper. In practice, it is where clean taxonomies meet messy reality. A name comes in from a spreadsheet, CRM export, metadata feed, or content archive. It looks obvious until you discover there are three people with the same name, two organizations that changed names after a merger, and a place entry that could refer to a city, a district, or a historical region. The difference between a useful link and a bad one often comes down to whether the evidence is visible, bounded, and understandable by the person reviewing the match.

That is why the design of the Wikidata + Google Knowledge Graph MCP server is interesting. It does not promise magical certainty. It focuses on a narrower, more disciplined job: helping an agent search Wikidata, inspect selected facts, and link local records to Wikidata QIDs with explicit evidence and explicit uncertainty when the evidence is not strong enough. That framing matters. In entity resolution, confidence without auditability is usually a liability.

The project, published as an open-source MCP server and CLI under an MIT license, sits in a very practical lane. It is read-only. It does not edit Wikidata. It does not edit Google. It does not edit user data. It also is not official software from Wikimedia or Google, and it is not an export of the Google Knowledge Graph. Those may sound like disclaimers, but they are also important operational boundaries. Good resolution tooling should make fewer claims than the people selling “single source of truth” platforms. The work is hard enough without pretending a lookup service can eliminate ambiguity.

Why evidence-first beats result-first

Most linking mistakes happen before the final match decision. They happen when too much candidate data arrives too quickly, often in a raw, unranked, or weakly explained form. Reviewers start pattern-matching on labels. Agents do the same if the prompt is not carefully constrained. A result list of fifty entities may feel powerful, but in practice it encourages casual overmatching.

This server takes the opposite approach. Search is bounded by default. It returns three candidates unless you explicitly ask for more, and even then the documented cap is five. That limit is not a weakness. It is one of the best design choices in the project.

When I have reviewed record resolution workflows, the biggest quality gains usually came from reducing noise, not adding more data. A reviewer deciding between three plausible candidates can inspect them. A reviewer facing thirty tends to anchor on the first promising label and move on. The same is true for agent behavior. Bounded search forces a more serious comparison between candidate entities and the local record.

There is also a second discipline at work here: selected-fact retrieval. Instead of pulling a broad, shapeless entity dump, the server supports reading selected facts, and can include ranks, qualifiers, and references on request. That combination is exactly what many resolution tasks need. You do not always need every statement about an entity. You need the statements that discriminate one candidate from another.

For a person named “John Williams,” a profession statement alone may not help. A birth date, field of work, nationality, and a ranked office held statement might. For an organization, the official website, industry, headquarters location, and inception date can matter more than a long description. Qualifiers and references can become decisive when two candidates are close and one has better-supported claims.

What the MCP setup actually enables

At a practical level, the server is designed to work with MCP clients such as Claude Code, Cursor, and Codex. If your workflow already lives in one of those environments, that matters more than flashy architecture diagrams. The point of MCP here is not abstraction for its own sake. It is standard access to tools that an agent can call in a predictable way.

The documented tools are straightforward:

  • kg_search
  • kg_entity
  • kg_related
  • kg_resolve
  • kg_status

That toolset covers the most common resolution loop. Search for candidates, inspect an entity, look at related context if needed, ask the resolver to evaluate a local record against candidates, and check system status. The CLI extends that into batch work and evidence export, which is often the difference between a promising demo and something teams can actually operate.

A quiet advantage of this setup is that Wikidata itself requires no account or API key for the documented use case, while the Google Knowledge Graph Search API is optional. That lowers friction for teams evaluating the workflow. You can begin with a pure Wikidata path and add the Google cross-check later if it truly improves your process.

The heart of the matter: explicit outcomes

One of the strongest signs that a resolution tool was designed by someone who has dealt with real matching problems is the presence of clear non-success outcomes. This project documents deterministic resolution logic with explicit results such as AUTO_MATCH, HOLD, AMBIGUOUS, and NO_CANDIDATE.

That vocabulary is healthy.

A lot of systems compress everything into “match” or “no match,” then bury edge cases inside threshold tuning. Here, the edge cases are first-class outcomes. If the evidence supports an automatic match, say so. If it does not, hold it. If there are multiple plausible targets, label it ambiguous. If nothing good appears, say there is no candidate.

This may seem obvious, but operationally it changes behavior. Teams become more willing to automate when they know uncertainty will be surfaced instead of smoothed over. Reviewers trust the system more when “I do not know” is an allowed answer.

I have seen too many pipelines where ambiguity is treated as failure and therefore engineered away through aggressive fallback logic. The result is often a deceptively high match rate paired with a slow accumulation of bad links. Those links are expensive. They contaminate downstream analytics, pollute recommendation layers, and create cleanup work that no one budgets for.

By contrast, deterministic logic and explicit outcomes create a review queue you can reason about. You can measure how many records land in HOLD versus AMBIGUOUS. You can examine whether certain source systems or record types generate more uncertainty. You can improve inputs rather than merely adjusting confidence thresholds.

What “evidence-first” looks like in practice

Suppose you are linking a local people directory to Wikidata. One local record contains a name, role title, employer, and maybe a city. Another contains only a name and a date. The first record is a strong candidate for structured resolution because multiple independent facts can be compared. The second may remain ambiguous even if the name is distinctive.

With a tool like kg_search, you start with a bounded set of candidates. Then, using kg_entity, you ask for selected facts that actually discriminate identity. If needed, you include ranks, qualifiers, and references. That last step deserves more attention than it usually gets.

Ranks help when a property has preferred and deprecated forms or multiple statements with different standing. Qualifiers help when a fact is true only in a certain period or context. References help when you need to inspect support rather than simply accept a statement at face value. In record resolution, those are not academic details. They often determine whether the evidence is strong enough to automate.

Consider a local organization record with a current name that used to belong to a different legal entity before a merger. A naked label match could be dangerous. A statement with qualifiers or chronology can reveal whether the candidate aligns with the present organization or a predecessor. The same goes for places whose names shifted over time, or public figures whose offices changed across terms.

The important thing is that the tool does not force you into bulk retrieval of every statement. That is good discipline. Evidence-first resolution is not about collecting more facts. It is about collecting the right facts, then retaining enough context to explain the decision.

The Google cross-check, used with restraint

The project also documents an optional Google cross-check through exact identifier joins. Specifically, it uses /m/ IDs matched to Wikidata property P646 and /g/ IDs matched to P2671. This is an elegant choice because it avoids vague “looks similar” reconciliation across providers. It is checking concordance through explicit identifiers.

Just as important, the project is careful about what that agreement means. Google and Wikidata alignment is treated as provider concordance, not proof of identity.

That distinction is one of the healthiest parts of the design. Many teams are tempted to think that two external systems agreeing must settle the matter. It often does not. It tells you that two providers connect the same identifiers or represent the same thing in a mutually consistent way. That is useful evidence. It is not the same as proving that your local record is definitely that entity.

This is where the combined phrase MCP for google knowledge graph and wikidata actually earns its keep. If you are evaluating MCP for google knowledge graph and wikidata together, the right question is not whether two knowledge sources can be blended. The right question is whether each source adds inspectable evidence without creating false certainty. On that measure, exact ID joins are much safer than semantic hand-waving.

In practical review workflows, I would treat the Google cross-check as a corroborating signal. Helpful, sometimes persuasive, never magical. If the local record is weak, cross-provider agreement should not overrule missing or contradictory core facts.

Where this fits among broader Wikidata MCP work

Wikidata itself documents a broader MCP capability for standardized programmatic exploration and querying via the Wikidata API and Wikidata Query Service. That broader context matters because it shows this project is not trying to replace all Wikidata access. It is solving a narrower and very operational problem inside that ecosystem.

The difference is useful. General exploration tools are great when you want breadth, open-ended querying, or exploratory analysis. Resolution tools need sharper edges. They need bounded search, selected evidence retrieval, and deterministic outcomes. Those requirements are easy to underappreciate until you try to move from “I can look up entities” to “I can safely link records at scale.”

If someone asks whether this is simply another MCP for Wikidata, the short answer is no, not exactly. It is an MCP for Wikidata designed with a specific resolution posture. That posture is what makes it valuable. General-purpose querying can tell you a lot about a graph. It does not automatically give you a disciplined framework for deciding identity.

What I would watch in a real implementation

Teams often assume the hard part is integration. Usually the harder part is deciding what evidence is sufficient for your domain. A music catalog, a nonprofit registry, a newsroom archive, and a biomedical reference set will all draw the line in different places.

The tool gives you explicit resolution outcomes, but your workflow still has to decide when AUTO_MATCH is acceptable without human review. In some settings, a mistaken link is mildly annoying. In others, it is legally or reputationally costly. The same evidence package that is acceptable for one domain may be too thin for another.

A sensible implementation usually starts by defining a narrow class of records for automation and a wider class for assisted review. For example:

  • Auto-match only when multiple high-value facts align and no strong contradictions appear
  • Route sparse records into HOLD even if the top candidate looks plausible
  • Treat name-only matches as suspect unless the entity is unusually distinctive
  • Use Google concordance as supporting evidence, not the deciding factor
  • Export evidence for audit from the first week, not after the first dispute

That is one of the few places where a list really helps, because these are operating rules rather than narrative ideas.

I would also pay attention to the human side of review. When evidence is exported, reviewers need to see enough context to understand why a match was proposed. If they have to reconstruct the decision from fragments, throughput collapses. The CLI’s evidence-export capability is therefore more important than it may appear. It turns the resolver from a black box into a reviewable process.

Why bounded search changes reviewer behavior

I want to come back to the three-candidate default, because it reflects a mature understanding of how people and agents behave under uncertainty.

When a resolver returns too many options, users often stop evaluating identity and start optimizing for speed. They look for a familiar label, a matching keyword, or a prominent description. Bounded search interrupts that reflex. It forces comparison. It also encourages better upstream query formulation because you cannot rely on drowning the user in possibilities.

There is a judgment trade-off here. A tighter result set can miss a relevant candidate if the search query is poor or the local record is noisy. But that is a visible problem. An overly broad result set creates invisible problems because low-quality matches slide through unnoticed. In real systems, visible misses are easier to remediate than hidden false positives.

This is why I generally prefer a workflow that accepts a modest number of NO_CANDIDATE outcomes over one that pretends everything has a likely target. If the input record is weak, uncertainty is the honest result.

The read-only boundary is a feature, not a limitation

Some teams instinctively want write-back capabilities in the same tool that performs matching. I think the read-only stance here is wise. Resolution and curation are related but different acts. Conflating them can create a false sense of authority, especially when an automated process appears to “improve” public knowledge bases.

By remaining read-only, the server keeps the epistemic boundary clean. It helps discover, inspect, and align. It does not claim the power to correct the underlying graph. That separation is healthy for governance. It also means teams can adopt the tool without entangling themselves in publication policies, edit review practices, or accidental data stewardship issues.

For organizations that need to maintain a local authority layer, this is often the better setup Wikidata MCP item anyway. You can store a proposed or accepted Wikidata QID in your own system, retain exported evidence, and decide separately whether any public knowledge base should later be updated by a human editor through the appropriate channels.

A measured view of the keyword hype

It is easy to see people search for phrases like MCP for google knowledge graph, MCP for Wikidata, or even the awkward combined keyword MCP for google knowledge graph and wikidata. Those searches usually signal a practical need: teams want agent-friendly access to structured entity data for search, enrichment, or linking.

The catch is that these phrases can flatten very different use cases into one bucket. An MCP for google knowledge graph might be useful for fast external cross-checks. An MCP for Wikidata might be useful for graph exploration or entity lookup. What distinguishes this project is not just that it touches both worlds. It is that it frames the interaction around evidence-first resolution and explicit uncertainty.

That is a stronger story than generic “knowledge graph integration.” It acknowledges that the hardest part is rarely fetching a candidate entity. The hard part is deciding whether the evidence supports identity, and documenting that decision in a way someone else can inspect later.

Where this approach is strongest, and where it is not

This server looks especially well suited to workflows where transparency matters more than raw volume. Editorial operations, metadata teams, research workflows, and authority control tasks all benefit from inspectable evidence and deterministic outcomes. The CLI and batch support suggest it can participate in larger processing runs, but the design still reads as quality-first rather than throughput-first.

That is not criticism. It is clarity.

If your ambition is to sweep millions of low-value records into approximate matches with minimal review, you would likely want a different system or a larger stack around this one. But if your aim is to build a record resolution workflow that an analyst, librarian, researcher, or data steward can defend, this design makes more sense.

The fact retrieval model also implies a discipline on the user side. You need to know which facts matter for the entities you are linking. A poor evidence schema will still produce poor decisions. The tool can expose ranks, qualifiers, and references, but it cannot decide for you which combinations are adequate for your governance standards.

The practical lesson

The most encouraging thing about this project is its modesty. It does not claim to solve identity universally. It narrows the task to something agents and humans can handle responsibly: search a small candidate set, inspect selected evidence, resolve deterministically when justified, and say “not enough” when the record does not support a clean decision.

That pattern is more valuable than many grander promises in the entity resolution market. Good linking systems are built as much on refusal and restraint as on successful matches. A resolver that can hold back is usually safer than one that is always eager.

If you are evaluating an MCP for Wikidata in a real operational setting, that is the benchmark I would use. Not how many entities it can surface, but how well it supports evidence, how honestly it represents uncertainty, and how easily a reviewer can understand why a QID was proposed. On those points, the design choices documented for this server are pointed in the right direction.