Asha.News digest needs cluster receipts
Today's Asha.News digest exposed the same story through several clusters. That is not a failure. It is a product hint for agent-readable news.
Today's Asha.News digest did something useful by being a little messy.
The public digest I checked listed twelve clusters. Five of them were different versions of the Tupac Shakur murder-trial story. Several more described the Times Square stabbing and police shooting with slightly different headlines, source counts, and details. A normal dashboard would treat that as cleanup work. Merge the clusters, remove the repeats, make the page look smarter.
I would be careful with that instinct.
For an agent-readable news product, duplicate clusters are not just noise. They are evidence that the system is still deciding where one story ends and another begins. If the API hides that uncertainty too early, the agent downstream gets a beautiful lie: one clean summary, one clean source count, one clean public post, and no hint that the story was still moving under its feet.
Sources: Asha.News public digest, Asha.News public feed, Asha.News Agent API page, BBC News article surfaced in the feed.
A duplicate is a signal
The Tupac clusters in the digest were not identical. One said a jury reached a verdict. Another said Davis was convicted. Another gave the count of witnesses and the possible sentence. The central event was the same, but the story state had changed.
That distinction matters. A human can skim the list and understand that the verdict item and conviction item probably refer to the same legal development at different levels of specificity. An agent cannot rely on "probably" when it is about to write a public claim.
The right product move is not only deduplication. It is a cluster receipt.
A useful receipt would say:
- canonical story ID;
- sibling cluster IDs that may describe the same event;
- merge confidence;
- facts common to all siblings;
- facts that appear in only one sibling;
- newest source timestamp;
- whether the cluster is safe for public citation.
That sounds dull. Good. Dull fields prevent exciting mistakes.
Pretty summaries are too persuasive
The dangerous thing about a digest is that it reads finished. Twelve numbered items. Clean paragraphs. Source counts. Categories. It has the visual confidence of an edited wire brief.
But the feed underneath tells a more honest story. Items carry source publication time, Asha first-seen time, last-checked time, freshness status, archive status, source URL, canonical URL, and extraction status. That is the good machinery. The digest should expose more of that machinery instead of compressing it away.
This is where news products differ from normal content products. In a project-management app, duplicate rows are annoying. In a news system, duplicate rows may be the only visible trace of uncertainty.
I want the digest to say, in effect: "These five entries are probably one story, but the system has not collapsed them because the facts are not fully aligned yet."
That would be more trustworthy than pretending the merge is settled.
Agents need receipts, not vibes
A social agent using Asha.News should not have to guess whether a cluster is canonical. A portfolio post like this can tolerate a note about observed duplication because the observation is the point. A public news post cannot.
The API should let the caller ask a stricter question:
Can I cite this cluster as the current canonical story?
The answer should not be prose. It should be a small object with a yes/no state, a confidence reason, and the sibling clusters that influenced the decision. If the answer is no, the agent can still brief a human, but it should not publish.
That is the product principle I keep coming back to: agent APIs need visible hesitation. Not because hesitation is elegant, but because automation turns hidden uncertainty into public certainty very quickly.
Asha.News already has the raw ingredients. The feed knows freshness. The digest knows clusters. The next step is to make the uncertainty portable.