Writing · August 2026
The distributed question.
Sometime this quarter, someone in your leadership meeting will ask a question like "are we still good to launch in Germany next week?" and the room will go quiet. Not because nobody knows. The legal review is in an email thread. The localization status is in a doc. The open bugs are in the tracker, the feature flag is in the code, support readiness is in a channel, and the commitment to the customer who was promised the date is in the notes from a sales call. Every tool involved is working exactly as designed, and the answer will still take two days and a volunteer with a spreadsheet. We have learned to call this a search problem. I do not think it is one, and the misdiagnosis matters, because better search is the fix the industry is building.
The questions that actually run a company are distributed questions: questions whose evidence is spread across tools, formats, permissions, and people's heads, and which no single system was ever asked to hold. The current fix, a model connected to all your tools, fails them in a specific way. It produces plausible answers to questions that need defensible ones.
The scale of the problem is not in dispute, though the precise numbers are. Vendor surveys converge on the same picture from different angles: Atlassian found teams spend about a quarter of their time searching for answers, Miro found roughly two thirds of knowledge workers say information is spread across too many tools, and one survey found about half hit at least one question a day where they do not even know who could answer it. Each vendor discovered the disease its product treats, so I trust the convergence more than any individual figure. The interpretation built on top of them is mine. It rests on two assumptions: that model capability keeps improving, and that the workspace record does not repair itself. And it has a falsifier I will return to: if wiring a frontier model into every tool starts producing answers teams can check and stand behind, the argument for changing the environment collapses into an argument for waiting.
Look closely at why the connected model fails, because the failure is more interesting than "hallucination." To answer the Germany question defensibly, a system must know which of three contradicting documents is authoritative, and "most recently modified" is not the same fact as "most recently true." It must know whether Thursday's thread superseded the January spec. It must recognize that the ticket in one room and the complaint in another are the same defect wearing different names. It must distinguish a decision from a musing that merely mentioned the same nouns, and it must know who actually owns the launch, which the org chart will not tell it. None of these are reasoning failures. They are missing facts. The tools that produced the record were built to transmit messages and store artifacts, not to preserve authority, supersession, ownership, or intent, so the record simply does not contain what the model is being asked to retrieve. A model cannot recover a decision that was never recorded as one. The bottleneck is not reading comprehension; it is that the workspace never wrote down what it knew.
Bolt a model on anyway and you get something genuinely useful that fails at the last step. It returns a confident, well-structured synthesis, and then a person has to verify it, and verification requires exactly the cross-tool crawl the answer was supposed to replace. Glean's own research is unusually honest about this: nearly nine in ten digital workers now use AI and most say it makes them personally more productive, yet barely one in eight say their organization performs significantly better, and workers report losing most of a working day each week to feeding AI context, checking its output, and cleaning up after its mistakes. Two in three admit to shipping AI work they have not fully verified. That last number is the quiet one, and I think it is the important one. A plausible answer does not remove the cost of a distributed question. It moves the cost to after the decision, where it compounds.
This is the problem Papaya is organized around, and I want to describe how we think about it rather than list features, because the thinking is the part I am confident will survive contact with reality. The bar we hold is defensibility: an answer should be one a person could repeat in a meeting with the evidence in hand. In our product reviews, defensibility works as a gate, not a score. An inference that cannot cite the reason it fired does not ship, however clever it is, because confident output the user cannot check is precisely what destroys trust in a system like this. That standard produces design language that sounds almost timid next to the industry default. "This may already be answered; here is what Alex wrote three weeks ago; judge for yourself," rather than "here is the answer." "No reply seen in this thread," rather than "you never replied," because the reply may have happened in a hallway. The system should claim exactly what it can see, attach what it saw, and say plainly where its sight ends.
The second premise runs deeper, and I hold it more strongly than the first: you cannot cite evidence that was never captured. If the durable failure is a record with the provenance missing, the durable fix is upstream of retrieval. When people and agents actually work in one shared place, the facts a model cannot infer, who decided, what superseded what, who owns which thing, get recorded as a byproduct of the work itself instead of reconstructed afterward by archaeology. That is why Papaya is a collaboration platform rather than a search index bolted onto everything else. Search reads the record. A workspace writes it.
The strongest objection is that I am describing the weaknesses of 2026 models and calling them permanent. It deserves its best form: context windows are growing, reasoning is improving, and several failures on my list will genuinely yield to capability. Multi-hop synthesis, reconciling aliases, representing disagreement instead of averaging it away: I expect frontier models to get markedly better at all three, and anyone building in this space should assume they will. But split the list. Authority, supersession, and intent are not reasoning problems. They are facts about what a group of people meant, and where the record does not contain them, a model's options are to guess or to ask. Guessing is how you get plausible. Asking, at scale, is how you get those lost hours back on someone else's calendar.
The second objection says the incumbent suites and the enterprise search vendors will simply build this, since they hold the data and the distribution. They might, and the gravity of existing data deserves respect; it has decided markets before. But the archive they hold has the same hole in it, because it was produced by decades of software optimized for message velocity, and adding provenance to those surfaces is not a feature so much as a different contract with the user: slower to write, more explicit about who decided what, honest about uncertainty. Rewriting that contract underneath hundreds of millions of users is the kind of thing incumbents do slowly and reluctantly, which is most of the opening a new workspace has. I could be wrong about their willingness. Whoever builds it, the contract is the part I am sure about.
To separate the strands: the descriptive claim is that teams are drowning in traces of work while the questions that run the company stay expensive, and that individual AI adoption has so far not closed that gap. My prediction, with more confidence in the direction than the timing, is that defensibility becomes the buying criterion for workplace AI, because the cost of unverified answers surfaces slowly and then all at once, in a shipped mistake with a paper trail leading back to a chatbot. The recommendation costs nothing and requires no belief in Papaya: treat any uncited answer, from a person or a machine, as a draft; ask "how would we check this" before asking "what does it say"; and move the work that matters into places that record it, because every quarter of unrecorded decisions makes every future question harder.
An organization does not run on answers. It runs on answers people can stand behind.
Colleagues taking each other's word is the oldest coordination technology there is, and it works because a person who says "I checked" can be asked how. As more of what we hand each other comes from machines, the choice is not between trusting AI and distrusting it. It is between systems that let people keep vouching for their work and systems that quietly make vouching impossible. The reason to insist on evidence is not suspicion of the machines. It is respect for the people who have to stand behind what the machines produce.
Papaya is the agent-native collaboration platform: a shared place for people and agents to communicate, coordinate, and get work done together. The longer argument is in the manifesto.
Sign up