How AI assistants work: architectures and tradeoffs | AI Wars of 2026
Technical drill-down for the assistant landscape. Source-backed analysis from AI Wars of 2026.
Loop Engineering
From architecture to practice
Loop Engineering: A Better Way to Think, Create, and Work with AI
Explore the practical side of working with AI. The companion site includes chapter resources, worked examples, prompts, and vendor comparisons.
As an Amazon Associate I earn from qualifying purchases.
How AI assistants work: architectures and tradeoffs
This chapter is the technical companion to AI Assistant Competitive Landscape.
Executive summary
AI assistants share components, but they do not all organize work in the same way: producing a grounded answer, scheduling constrained work, investigating a question, and executing an open-ended task create different requirements for context, control, and verification. The documented patterns below provide a more useful basis for comparison than a universal component-count score. One product can support several patterns, and multi-agent coordination can wrap a research or action workflow rather than form a separate class of product. Choose an architecture for its task fit, then test correct authorized completion, human effort, cost, and recovery; the sources reviewed do not establish an overall performance ranking. C2 C9 C12 X4 E5 E6
Introduction
An assistant that answers a policy question and an assistant that changes a customer record may use similar models but do different work. The first needs relevant, authorized evidence and a supported answer. The second also needs permission to change the record, a valid operation, confirmation of the resulting state, and a response to partial failure. Counting both systems' models, connectors, or memories does not establish which is better.
Here, architecture means the arrangement of context, decisions, tools, state, and oversight around a task. It does not mean the identity of the underlying LLM. We group product modes and configured workflows, not entire companies. Microsoft 365 Copilot's documented answer flow is one example; it does not describe every Copilot product. A research mode and an action mode from the same provider may need different diagrams.
The grouping has direct support in engineering descriptions. Anthropic distinguishes predefined workflows from dynamically directed agents and describes patterns that can be combined. Its Research engineering account documents a coordinator delegating independent investigations to workers. Those sources support recurring patterns, not exclusive market clusters or a ladder on which more elaborate designs are always stronger. The older pattern article now warns that its tooling examples have changed; we use its design distinctions, not those examples as a current product inventory. E5 E6
Each section leads with a conceptual architecture diagram, then identifies assistant modes with supporting documentation and the potential advantages and disadvantages of that arrangement. The diagrams are abstractions, not reverse-engineered production blueprints. Dashed connections are optional or configuration-dependent. In particular, an outcome-checking box identifies what must be evaluated; it does not certify that every named product performs an independent check.
The pros and cons are architectural tradeoffs to test, not measured product results. Primary documentation verifies what a provider describes, while independent evaluation is still needed to establish how well it works. Detailed product evidence, product-scope limits, and a proposed testing method follow in the appendices. Source notes identify the documentation behind the descriptions; no overall strength score is assigned.
Architecture 1: Conversational and grounded-answer assistants
How it works. The main deliverable is an answer, explanation, conversation, or draft. Retrieval and memory can improve the context supplied to the model, but the task does not require the assistant to commit an external action. This functional view does not imply a single model call or disclose the provider's internal routing.
Assistants that work in a similar fashion
| Assistant mode | Evidence for the fit | Important qualification |
|---|---|---|
| Microsoft 365 Copilot's documented question-answering flow | Microsoft 365 prompt, permission-scoped Graph grounding, LLM response and returned output. C9 | Not a complete diagram of its app actions or Copilot Studio. |
| Gemini in Workspace and consumer Copilot's connected-source retrieval | Authorized organizational or personal context contributes to responses. C6 C8 | Keep personal and organizational accounts separate; history is not automatically semantic memory. |
| Perplexity Search and source-configured You.com custom agents | Web/files/source selection and generated answers are described. D10 X13 | Deep Research and browser execution are different modes. |
| Grok's documented web/public-X answering | Search and supplied media/documents inform conversation. D3 | Its complete runtime and connector action scope are not established here. |
| Pi's memory-supported conversation | Cross-chat preferences and explicit memory management are documented. X1 | Conversational fit, not proof of the exact inference path shown above or of external tools. |
| Selected conversational bots on Poe | Bot choice, chat context and optional Memory are documented. X2 | Classify the selected bot; Poe as a whole is a platform containing different capabilities. |
Pros and cons
| Potential advantages | Tradeoffs and failure modes |
|---|---|
| Appropriate for explanation, drafting, and questions where the person remains the executor. | The user still has to turn the answer into an external result; a polished response is not completed work. |
| Relevant retrieval can supply facts missing from the model's prior knowledge. | Missing, stale, unauthorized or misleading context can undermine the answer. Citations can be present without supporting the claims. |
| Personal memory can reduce repeated explanation of preferences. | Incorrect or obsolete memories can carry into later conversations; correction and deletion need testing. |
| A narrowly scoped answer service can avoid unnecessary action permissions. | Read access still creates privacy and untrusted-content risks. “Read-only” is not equivalent to harmless. |
Best fit: conversation, grounded questions, explanation, and drafts reviewed by a person.
What would establish strength: supported answers, relevant coverage, good handling of missing evidence, privacy-respecting context use, and lower human correction effort on matched tasks.
Architecture 2: Constrained workflows and schedulers
How it works. A predefined process or constrained optimization problem determines much of the work. An LLM can help interpret inputs or perform a step without directing the whole workflow. A scheduler can adapt to changed deadlines without being a general-purpose autonomous agent. E5
Assistants that work in a similar fashion
| Assistant mode | Evidence for the fit | Important qualification |
|---|---|---|
| Reclaim 1.0 scheduling | Priority-, deadline-, availability- and rule-based placement/rescheduling are described. X4 | Its event-protection rules matter; do not assume the same behavior in Reclaim 2.0. |
| Motion's automatic task planning | Tasks, priorities, deadlines and available time drive planning and replanning. X6 | This classifies scheduling behavior, not all Motion chat or meeting features. The solver's internal implementation is not established. |
| Configured Bardeen automations | The support overview describes creating and running automations in the browser or cloud. X8 | A candidate fit for defined automations; the overview does not prove every current workflow uses a fixed control path. |
| Reclaim 2.0 previewed calendar changes | Preview and review/apply controls are documented. X5 | A private-beta hybrid, not the same automatic scheduling behavior as 1.0; suggested tasks must not be described as automatically scheduled. |
Pros and cons
| Potential advantages | Tradeoffs and failure modes |
|---|---|
| Explicit rules make expected behavior and failure cases easier to specify. | A workflow can execute the wrong business rule consistently; predictability is not correctness. |
| A narrow action space can reduce unnecessary exploration and calls. | Novel exceptions may fall outside the process and require redesign or human intervention. |
| Constraints and intermediate outputs provide concrete places to check the result. | Incorrect priorities, missing data or incompatible constraints can produce an unwanted result or no feasible plan. |
| Replanning can keep schedules aligned with changing work. | Frequent changes can create calendar churn or interfere with commitments; these costs must be measured. |
Best fit: repeatable, bounded operations with explicit constraints and observable success conditions.
What would establish strength: feasible results, faithful constraint handling, fewer unwanted changes, and less human effort than manual work or a simpler automation.
Architecture 3: Adaptive research agents
How it works. The system chooses subsequent searches in response to what it finds. Its principal outcome is an evidence-based artifact rather than a change to a customer account, calendar, or production system. This is an adaptive-agent specialization, not an entirely different technological foundation from the action architecture. “Assess evidence” is a function to evaluate, not a claim that the assistant independently verifies every source.
Assistants that work in a similar fashion
| Assistant mode | Evidence for the fit | Important qualification |
|---|---|---|
| ChatGPT deep research | An editable plan, selected sources, progress/interruption, and cited results are documented. Connected-app operations are read-only. C2 | Do not transfer ChatGPT Work's write capabilities into this mode. |
| Claude Research | Anthropic describes a lead researcher and parallel workers that adapt their searches. E6 | It combines this pattern with the coordination architecture below. The engineering account is not an independent performance comparison. |
| Perplexity Deep Research | Official help identifies the research mode and source-based synthesis. D10 | The bounded source describes the mode, not the full internal loop; use a provisional architectural assignment. |
| You.com advanced research and research tasks in Mistral Vibe Work | Source controls and research/reasoning settings or stepwise tool work are described. X10 X13 | Task-level candidates, not proof that their implementation matches this diagram. Vibe Code belongs to a different execution setting. |
Pros and cons
| Potential advantages | Tradeoffs and failure modes |
|---|---|
| Iterative searching can address questions whose useful sources are not known beforehand. | The system can follow weak leads, repeat searches, or stop before finding material counterevidence. |
| Plans, citations and intermediate findings can make the investigation inspectable. | An inspectable process can still rely on unsupported claims, inaccessible sources, or circular citations. |
| Read-only source use can keep the workflow separate from consequential writes. | Retrieved material remains an untrusted input and may contain sensitive or misleading information. |
| Search effort can be adapted to the question's complexity. | More browsing can increase latency and cost without improving the answer; a defined stopping rule matters. |
Best fit: questions requiring multiple sources, uncertain search paths, and a reviewable synthesis.
What would establish strength: better claim support and material coverage, calibrated uncertainty, and acceptable effort/cost compared with simpler retrieval and human-led research.
Architecture 4: Adaptive action agents
How it works. The system selects actions, observes feedback, and changes its plan while pursuing a goal. Actions may alter external state, so authority, confirmation, duplicate prevention and recovery become central. A depicted result check is not proof of a product-wide verification guarantee. C12 E5
Assistants that work in a similar fashion
| Assistant mode | Evidence for the fit | Important qualification |
|---|---|---|
| Claude Code | A context-action-verification loop, tool access, sessions and permission/checkpoint mechanisms are described. C12 | File checkpoints do not reverse remote API, database or deployment effects. |
| ChatGPT Work | Browser/connected-app work, deliverables, continuing tasks, progress review and important-action approvals are documented. C3 | Local and cloud execution differ; complete recovery and tool-selection contracts remain unestablished. |
| Meta Muse | Vendor engineering describes its runtime, browser work, connector execution and separate permission decisions. D2 | VM restoration is not external transaction rollback; security effectiveness is not independently tested here. |
| Configured OpenClaw agents | The runtime describes context assembly, inference, tools, persistence and specified interruption/retry behavior. X9 | The operator's configuration determines actual models, tools and authority. The runtime name alone does not identify one deployed assistant. |
| Comet browser-action mode | Browser/tab context and website actions are documented. D11 D12 | The detailed planning loop is not disclosed in the selected pages; enterprise policy is distinct from consumer settings. |
| Mistral Vibe Work/Code and Lindy's cross-tool work | Stepwise tools and sensitive-action approval, or approved team write actions, are described. X3 X10 X11 | Similar functional aims, with less complete implementation evidence here. Scheduled routines can also follow the constrained-workflow pattern. |
Pros and cons
| Potential advantages | Tradeoffs and failure modes |
|---|---|
| Can pursue tasks whose exact steps depend on changing environment feedback. | Incorrect intermediate actions can compound; flexibility makes the path less predictable. |
| Can deliver a changed record, file or completed operation rather than advice alone. | Wrong writes can have effects that cannot be undone by resetting the assistant. |
| Can switch tactics when a tool fails or information is missing. | Retries can duplicate external actions, consume budget, or conceal partial completion unless the implementation handles them. |
| User approval and scoped tools can bound consequential actions. | Approval fatigue, ambiguous requests and misconfigured permissions can undermine those controls. |
Best fit: bounded goals requiring adaptable execution, where progress and destination state can be checked.
What would establish strength: correct authorized completion under normal and failure conditions, with manageable intervention, bounded cost, and explicit recovery limits.
Architecture 5: Coordinated multi-agent systems
How it works. A coordinator divides work among agents with their own instructions, context or tools, then integrates the results. The worker count shown is illustrative. Workers may themselves be research agents or action agents, so this is a composition pattern across the other families, not a higher capability tier. Model routing, multiple API calls, or a list of specialist functions does not alone establish multi-agent coordination. E5 E6
Assistants that work in a similar fashion
| Assistant mode | Evidence for the fit | Important qualification |
|---|---|---|
| Anthropic's documented Claude Research system | A lead agent develops a research strategy and delegates parallel searches to workers. E6 | Vendor engineering account and internal evaluation, not a matched independent comparison with other assistants. |
| Muse's documented subagent-assisted execution | The engineering account describes concurrent subagents and a specialized browser subagent. D2 | The full product need not use parallel workers for every request, and the browser/security boundary must not be flattened into the generic drawing. |
| Claude Code with configured subagents | Subagents are described within the Code tool/runtime documentation. C12 | Delegation-enabled mode, not a claim that every Code session runs multiple agents. |
| Developer implementations using agent SDKs | Delegation, handoffs and subagent facilities are described. C4 C13 | Configurable building blocks, not evidence that the corresponding consumer assistant exposes or uses this exact architecture. |
Pros and cons
| Potential advantages | Tradeoffs and failure modes |
|---|---|
| Independent investigations can run in parallel with separate context. | Tightly dependent tasks may not parallelize; coordination can add more work than it saves. |
| Workers can specialize in a bounded question or tool environment. | Poor delegation creates duplicated searches, missing coverage or incompatible outputs. |
| Separate contexts can reduce interference among unrelated subtasks. | Summaries can omit evidence or constraints needed by the coordinator; inconsistent state can persist. |
| A coordinator can reconcile several perspectives before returning a result. | Agreement is not independence or correctness. More workers can increase token use, latency, and correlated error. |
Anthropic reports gains for selected breadth-first research tasks, alongside substantial resource costs and limitations on tasks requiring heavily shared context. That supports a conditional design choice, not “multi-agent is strongest.” E6
Best fit: valuable work with genuinely separable subtasks and an effective integration step.
What would establish strength: better task outcomes or useful time savings than a strong single-agent baseline under declared resource budgets, after accounting for coordination and human review.
Choosing an architecture without forcing every assistant into one group
| If the task primarily requires... | Start by evaluating... | Do not assume... |
|---|---|---|
| A useful answer or draft | Conversational/grounded-answer mode | That it needs write authority or a persistent autonomous process. |
| Repeated work under explicit rules | A constrained workflow or scheduler | That fixed processes are weaker than agent-driven ones. |
| Finding and synthesizing evidence along uncertain paths | An adaptive research agent | That more sources or citations imply a better-supported result. |
| Changing external state through an uncertain sequence | An adaptive action agent | That permission to read implies permission to commit. |
| Several separable investigations or work packages | A coordinator-worker arrangement | That splitting work always improves it. |
Keep placement and access dimensions separate. Browser versus app, device versus cloud, personal versus enterprise, persistent versus session-only, and one model versus model routing describe additional dimensions. They can be drawn as labeled boundaries around several of the patterns above.
- Alexa+ and Siri/Apple Intelligence combine conversational, routing and action functions across service or device boundaries. Use task-specific views; the selected sources do not justify a single exclusive placement or a complete shared internal algorithm. D5 D6 D7 D8 D9
- Poe is a platform of bots; classify the selected bot and mode, not the catalog as if every feature belonged to one runtime. X2
- Phind remains unassigned because no current readable primary product body was obtained. Lack of access is not an architecture. X14
- Other provisional fits remain marked in the tables rather than being forced into an exact blueprint. A less detailed disclosure does not establish a weaker product.
The shared component vocabulary remains useful, but as a comparison aid across these diagrams rather than a universal machine that every product must resemble. The proposed groups would need revision if product-level tracing showed materially different control paths, or if independent reviewers could not assign the same modes consistently. They have not been established as statistical market clusters.
Appendix A: Test the premise: does more mean stronger?
Competing hypotheses
H1 — component complementarity. A relevant new component improves a defined outcome when it removes a real limitation: authorized organizational retrieval may improve document-grounded answers; an action tool may enable a workflow previously limited to drafting; saved state may prevent loss of completed work.
H2 — coordination and exposure cost. Additional tools, memories, agents, or integrations introduce selection ambiguity, irrelevant context, stale state, latency, cost, and opportunities for unintended actions. A narrower system may perform better on its intended task.
Both have concrete support for investigation:
| Evidence | What it supports | What it does not establish |
|---|---|---|
| Copilot Studio documentation warns that overlapping topic descriptions can make selection unpredictable; its disambiguation behavior has stated limitations. C10 | Tool/topic selection quality matters; increasing the available set is not automatically helpful. | A measured penalty per added component, or the relative quality of consumer assistants. |
| Agentless describes a fixed localization, repair, and patch-validation process and reports competitive results on the SWE-bench Lite setting studied. E1 E2 | Include a simple workflow baseline before attributing gains to autonomous architecture. | That a simpler system always wins, or that its older coding results rank current assistants. This is not a clean commercial-product component-count experiment. |
| AgentDojo evaluates useful task completion and adversarial behavior in a stateful tool environment using checks on environment state. E3 | Evaluate utility and security together; access to useful tools also creates consequential failure paths. | That any assistant in this document has the reported benchmark behavior. |
| The original tau-bench repository combines user-agent dialogue, domain API tools, and policy guidelines, and warns that its tasks are outdated. E4 | Test policy-constrained multi-step behavior and pin the benchmark version and user simulator. | A current leaderboard for these products, or a guarantee that an older task set remains valid. |
Decision: use H1 to generate specific advantage hypotheses, and H2 to design the counter-tests. Reject “more components therefore stronger” as an unvalidated inference.
Would revise: a matched, repeated, independently checked evaluation shows a meaningful gain from a specified component change at an acceptable cost and risk. A larger product catalog, vendor testimonial, or model benchmark alone would not meet that condition.
Appendix B: Component model and evidence vocabulary
These are functional slots, not proprietary modules that every vendor must implement separately. A deterministic scheduler need not contain an autonomous planning loop. A read-only research assistant need not have authority to purchase anything.
| Slot | Function | What would make it better for a task? | Common false equivalence |
|---|---|---|---|
| Model and inference | Interpret, reason, generate, or select a proposed action. | Correct decisions for the task at suitable latency and cost. | A newer model name equals a stronger deployed assistant. |
| Grounding and retrieval | Supply relevant external or organizational evidence. | Authorized, current, relevant information with support for generated claims. | More indexed documents equal better answers. |
| Memory and task state | Retain preferences, conversation, plans, or execution progress. | Appropriate recall, freshness, correction, deletion, and safe resumption. | Chat history, semantic memory, and a durable workflow checkpoint are the same thing. |
| Orchestration | Choose and sequence work; route among tools or agents; handle continuation. | Correct routing, bounded execution, useful delegation, and predictable failure behavior. | More agents or a longer chain means better reasoning. |
| Tools and action interfaces | Read or change external state through APIs, browsers, devices, or code execution. | The necessary operation works on the required object with valid arguments. | A connector count measures action coverage, depth, or reliability. |
| Identity, permissions, and authority | Establish whose data/action is involved and whether execution is allowed. | Least privilege, understandable approvals, revocation, and policy enforcement. | Read permission or technical access authorizes a consequential commitment. |
| Observation, verification, and recovery | Expose activity, check the outcome, and deal with interruption or error. | Accurate receipts, external-state checks, safe retry, and explicit recovery limits. | A completion message proves success; local restore reverses external effects. |
| User interface and distribution | Give people a place to direct, inspect, interrupt, and receive the work. | Low-friction access with usable control and feedback. | Being a default app proves trust, retention, or useful delegation. |
Evidence states for diagrams
- D — documented mechanism: a primary technical/help page describes a particular behavior. Not independently tested here.
- A — announced or described: product copy or a high-level vendor explanation; implementation/availability can be limited.
- U — unestablished: not established by the sources checked. Not “absent.”
- F — future/beta qualification: retain the source's rollout or development boundary beside the affected capability.
- T — task-tested: reserved for a specified, reproducible evaluation. No product-wide T ratings are awarded in this document.
Do not turn D/A/U into 3/2/0 and add them. A proprietary assistant may disclose less than an open runtime while performing better. That is a documentation gap, not a performance score.
Use equal-sized slots. Color can indicate functional role, while letter labels indicate evidence state. A missing or undisclosed component receives an explicit U, not a blank cell that looks like zero.
Appendix C: Detailed product/component evidence map
The “possible advantage” column is a hypothesis to test, not a finding of superiority. References apply to the mechanisms in their row. Product eligibility, configuration, connected-account permission, and task choice remain part of the comparison unit.
General-purpose and workplace products
| Product scope | Documented components | Scope limits and unknowns | Possible advantage and decisive test |
|---|---|---|---|
| ChatGPT Work | Project chats/files/instructions, connected apps, browser work, deliverables, cross-surface continuation, scheduled/event-triggered work under eligibility controls; progress review and approval of important actions. C3 | Cloud Work, local Work, and Codex Local have distinct controls. Exact recovery guarantees and selection of persistent memories are U. The older “ChatGPT agent” article has conflicting availability wording. C1 | Multi-step work using the relevant apps; test completed external state, intervention burden, and continuity across interruptions. |
| ChatGPT deep research | Reviewable plan; web, files, selected sites and supported connected sources; progress/interruption; citations, source list and activity history. Connected-app use is read-only. C2 | Not a general write-action branch. App availability and organizational RBAC constrain access. Saved-memory behavior and automatic recovery are U in this source. | Auditable research; test claim support, omissions, source restrictions, and correction time against ordinary search/chat. |
| Gemini Apps, personal accounts | Connected-app retrieval/actions; explicit app selection or automatic selection; app-based personalization; custom MCP app connections described in current help. Users connect/disconnect accounts and permit data access. C5 | This help is for personal accounts, not Workspace. Device, location and activity settings affect access. Uniform action approval, execution traces, and rollback are U. | Tasks using connected personal context; test the actual allowed app operations and stale/unauthorized-context handling. |
| Gemini in Workspace | Authorized Workspace grounding; existing access controls, with encryption and information-rights restrictions; application-specific side-panel history and administrator retention controls. C6 | Workspace history is not one universal shared memory. The privacy hub does not establish a common executor/recovery design across every app. | Organizational questions under existing access controls; test evidence accuracy and denial cases against the same corpus and permissions. |
| Microsoft Copilot, consumer | Authorized retrieval from connected personal services; account connection and per-conversation connector controls; conversation history with connected answers. C8 | Concrete examples establish retrieval, not a general autonomous write system. Do not borrow Microsoft 365 tenant or Studio capabilities. Persistent-memory mechanism, action approvals, and recovery are U here. | Personal information finding; test retrieval relevance and permitted source access rather than assuming enterprise automation. |
| Microsoft 365 Copilot | Microsoft 365 app prompt, Microsoft Graph grounding, LLM response, and returned output; signed-in-user data access, Conditional Access/MFA and stored chat interactions. C9 | Tenant residence does not confer tenant-wide access. Chat history is not proof of separate semantic memory. App-specific action and rollback contracts are not established by this architecture overview. | Grounded work in a permissioned Microsoft 365 corpus; test citations, missing context, and denied records. No same-task advantage is measured here. |
| Claude consumer chat/desktop | Remote MCP connectors and desktop local extensions; retrieval and actions; per-account access; tools configurable as Always allow, Needs approval, or Blocked; reconnection controls. C11 | Team/Enterprise organization setup differs from individual connection. Connector docs do not establish a universal memory/research loop, action trace, or rollback. Claude Code is a different surface. | Tool-enabled work with explicit permission choices; test tool coverage, approval fidelity, blocked-action attempts, and correct destination state. |
Consumer, device, and browser products
| Product scope | Documented or described components | Scope limits and unknowns | Possible advantage and decisive test |
|---|---|---|---|
| Meta Muse | Vendor engineering describes a per-user cloud VM, Hatch agent runtime, memory files, browser subagent, connector workers and separate credential/security services. Sentinel evaluates actions/egress and returns allow/deny/ask; approvals bypass conversational assent. Audit trail, browser takeover and VM backups are described. D1 D2 | Inference can send limited data outside the VM. Confidential VM is future, not the launch protection. Restoring the VM does not reverse an external message, booking, or payment. Comparative security claims remain untested. | Persistent cross-app work with an explicit control plane; test permission enforcement, takeover, injection resistance and external-effect recovery separately. |
| Grok, consumer apps/web | Web and public-X retrieval, uploaded media/documents and conversation; public connector page describes searching/referencing connected tools; sharing, training and private-chat controls. D3 D4 | API/enterprise and Grok on X are separate scopes. Public connector page requires sign-in for further detail. Persistent memory, write tools, action approvals, orchestration and recovery are U. FAQ date and current branding differ. | Fresh-source conversational analysis; test source fidelity and temporal accuracy. No supported basis here for a write-action diagram. |
| Alexa+ | Amazon describes model routing, grounding, household/personal context and remembered preferences; task “experts,” API orchestration, website interaction and device/smart-home actions. D5 D6 | Updated availability overview and older launch architecture are not a current implementation audit. Uniform transaction approval, credential isolation and rollback are U. Preference confirmation is not purchase approval. | Household/device/service tasks; test actual device/service state, household identity, confirmations and completion under failure. |
| Siri / Apple Intelligence | Current feature page describes Siri AI context and app actions. Privacy documentation distinguishes on-device processing, PCC and off-device request logging. Optional ChatGPT extension has separate account/data-policy and consent behavior. D7 D8 D9 | Siri AI beta, older Siri behavior, supported hardware, region and OS version must be separated. ChatGPT settings are not identical across Siri modes. A PCC report is not an app-action success receipt. | Personal/device work with relevant local context; test eligible configurations, action correctness and data boundaries. Do not treat every Siri device as having the beta's features. |
| Perplexity search | Real-time web retrieval, synthesized answers, model/mode choice and source citations. D10 | Citation presence does not establish citation support or completeness. Persistent memory, account-action permissions and transaction recovery are U in this source set. Do not transfer Comet powers into search. | Source-led question answering; test groundedness, source quality, omissions and research effort. |
| Comet browser assistant | Tab/page context, visible-text/metadata extraction and website actions; parallel errands described. Enterprise policy supports browser control, case-by-case consent, Always Allow and domain rules. D11 D12 | Enterprise-admin controls are not consumer defaults. Domain rules can override global settings; effective policy must be tested. General postcondition verification and rollback are U. | Browser tasks on permitted sites; test domain-policy conflicts, actual form/transaction state, malicious page content and recovery. |
Specialists, aggregators, and operator-run systems
| Product scope | Documented or described components | Scope limits and unknowns | Possible advantage and decisive test |
|---|---|---|---|
| Pi | Cross-chat personal memory; explicit remember/forget requests; settings for adding, deleting and updating memory. Memory feedback does not itself delete an item. X1 | Tool execution, external grounding, integration inventory and transactional approvals/recovery are U. A memory instruction may need repeating, according to the help page. | Ongoing conversation and preferences; test useful recall, correction/deletion and inappropriate carryover. |
| Poe | Multi-provider/user-created bots; managed chat context; optional cross-chat Memory with exclusions; temporary chats; some bots can execute code or provide interactive canvases. X2 | Poe is not one homogeneous model/planner/toolset. Bot identity and Memory eligibility matter. Universal external-account approvals and recovery are U. | Model/bot selection for a task; compare named bots at pinned context settings rather than awarding Poe every bot's capability at once. |
| Lindy | Current overview describes Slack-centered team work, Skills, triggered/scheduled Routines, editable memory Files and meeting records. Cross-tool work and each person's credentials are described; write actions are stated to wait for approval. X3 | High-level documentation/security assertions, not an audit. Detailed permission propagation, approval exceptions, idempotency and rollback are U. | Repeatable team operations; test correct records, account separation, approval fidelity, duplicate prevention and partial completion. |
| Reclaim 1.0 | Priority-, deadline-, availability- and rule-based calendar placement/rescheduling; treatment of non-Reclaim events is explicit. X4 | A constrained scheduler, not evidence of a general autonomous assistant. User priorities change scheduling behavior. | Calendar optimization; test feasible placement, protected events, churn and missed deadlines against manual/rule-based baselines. |
| Reclaim 2.0 | Assistant, preferences/policies, background calendar agents, connected task context and MCP integration described. Preview Mode precedes Review & apply changes. X5 | Private beta. Tasks are described as suggested rather than automatically scheduled; a meeting feature is in development. Do not merge its behavior with 1.0. | Previewed schedule changes; test preview-to-commit fidelity and conflicts at the actual beta configuration. |
| Motion | Task/deadline/priority scheduling, automatic replanning, task creation/update through AI Chat, and meeting-to-task handling described in help. X6 | Detailed memory, connector permissions, approval gates, audit and rollback are U. Replanning is not a general recovery guarantee; do not copy Reclaim's controls into Motion. | Deadline-constrained task planning; test feasible schedules, faithful task extraction, correction effort and actual completion. |
| Bardeen | Web extraction/search, enrichment, qualification and exports; support describes natural-language automations and local-browser/cloud execution. X7 X8 | Mostly product/support overview. Detailed state, authorization, approvals and recovery are U. Legacy terminology is not proof of the current architecture. | Research/data-preparation workflows; test extraction accuracy, provenance, duplicates, export correctness and resistance to website changes. |
| OpenClaw | Technical docs specify intake, context assembly, inference, tools, streaming and persistence; queues, hooks, run ownership, transcripts, lifecycle/tool events, timeouts and specified retry/duplicate-suppression behavior. X9 | Operator-run configuration determines model, channels, permissions and exposure. Hooks are not proof of safe policy. Universal human approvals and business-action rollback are U. | Configurable stateful workflows; test stale writes, duplicate sends, interruption, authorization and external side effects in the actual installation. |
| Mistral Vibe | Current docs describe Work and Code, context from files/web/tools, stepwise work, visible tool calls, steer/stop and approval before sensitive actions. Code CLI setup is separately documented. X10 X11 | Do not assert “formerly Le Chat” from these pages. Work/Chat migration is described progressively; the “September 22” passage has no year. CLI model/provider and filesystem details belong to Code, not automatically Work. | Test Work research/drafting separately from Code repository work; evaluate groundedness, correct edits and permission behavior. |
| You.com custom agents, web UI | Model selection, files, controlled web/source access, personalization/settings, named agents and sharing controls. X12 X13 | General workflow/internal-tool integration language is high-level. Research API capabilities cannot automatically be attributed to the consumer UI. Write approvals, action logs and recovery are U. | Source-constrained research/reporting; test source restrictions, citation support, omissions and sharing boundaries. |
| Phind | No readable primary product body was obtained in this pass. X14 | Direct requests returned HTTP 403 and a browser attempt failed. Operational status and current components remain U. Access failure does not prove shutdown. | Establish first-party status before designing a component comparison or testing developer-search claims. |
Developer runtimes: useful evidence, separate panels
These are not additional capabilities to silently add to the corresponding consumer product rows:
- OpenAI Agents SDK: tools, loops, handoffs, guardrails, MCP, sessions and tracing are developer-runtime affordances. Direct Responses API use may leave loop, dispatch and state ownership with the developer. C4
- Gemini API: the model proposes function names/arguments; the application executes the function. Validation, API authentication and recovery remain implementation work. C7
- Copilot Studio: generative orchestration selects tools, topics, knowledge and agents. Testing exposes an activity map; reset/cancel behavior has specific limits. This is not the architecture of every Microsoft 365 or consumer Copilot experience. C10
- Claude Code / Agent SDK: context-action-verification loop, tools, sessions, permissions and checkpoints are documented. Remote database/API/deployment changes are not covered by file checkpoints. SDK affordances are not consumer chat defaults. C12 C13
Appendix D: Product identity and scope
These distinctions limit which product features can be compared.
| Existing shorthand or potential inference | Validated boundary |
|---|---|
| “ChatGPT agent” is a clean current product label. | The help page's Overview says it is no longer available while lower sections retain legacy operation/availability wording. Work has separate current documentation. Do not invent a retirement date or combine every section as current. C1 C3 |
| “Perplexity + Comet” is one component inventory. | Separate cited search answers from browser context/actions and enterprise policy. D10 D11 D12 |
| “Reclaim + Motion” is one scheduling system. | Separate vendors, and separate Reclaim 1.0 from 2.0 private beta. X4 X5 X6 |
| “Mistral Vibe, formerly Le Chat.” | Official pages establish current Work/Code modes and chat-endpoint continuity, but not that exact rebrand assertion. X10 X11 |
| Phind's inaccessible website proves closure. | It proves this retrieval failed. A primary shutdown statement was not obtained. X14 |
| Muse's VM is confidential computing today, and all inference stays there. | The engineering account distinguishes current VM controls from future Confidential VM work and describes inference data leaving the personal VM. D2 |
| Siri/PCC/ChatGPT share one privacy and consent boundary. | Draw separate device, PCC and third-party branches; feature/account/configuration conditions matter. D8 D9 |
Appendix E: Validation plan before any claim of strength
Define the comparison unit
Record product + mode + plan + model/version when disclosed + device/OS + region + connected accounts + permissions + settings + test date.
For hosted products that silently change, report the observed configuration and dates; do not pretend it is a reproducible pinned model. For open runtimes, record the runtime revision and operator configuration. An open runtime is not a directly comparable “assistant” until it is configured.
Begin with separate task families
| Task family | Example test objective | Success must be checked in |
|---|---|---|
| Grounded research | Answer a question using a permitted source set and identify material uncertainty. | Source passages and blinded factual/coverage review. |
| Permissioned workplace retrieval | Find the authorized policy or customer record without exposing a denied record. | Gold corpus, access-control expectations and retrieved evidence. |
| Calendar planning | Propose or apply a feasible schedule subject to protected events and priorities. | Calendar/task state, conflicts and user-specified constraints. |
| Browser/application action | Complete a form or update a sandbox record after the required approval. | Actual destination state and an action receipt, not the assistant's final message. |
| Coding work | Make a scoped repository change without regressions. | Held-out tests, patch review and repository state. |
| Personal memory | Use a preference, update it after correction, and stop using it after deletion. | Predeclared recall/correction/deletion checks over separate sessions. |
These are evaluation proposals, not claims that every listed product supports every task. Use synthetic or permissioned sandbox data and avoid live purchases, external messages, or destructive actions.
Proposed measures and their denominators
| Measure | Definition for this study | Boundary |
|---|---|---|
| Correct authorized completion | Trials satisfying both the outcome checks and required authority/policy conditions / all eligible assigned trials. | Authorized refusal is scored according to the task's expected policy outcome, not automatically as failure. |
| Intervention burden | Human active time, approvals, corrections, and takeovers per assigned trial and per successful trial. | Separate useful approval from rescue after failure. |
| Cost per successful task | Total declared trial cost, including failed attempts, / successful tasks. | State included charges and allocation assumptions; if there are no successes, the ratio is undefined. Subscription allocation is an assumption, not observed marginal cost. |
| Latency | End-to-end time per trial, with median and a tail statistic. | Report timeouts and incomplete runs; do not silently drop them. |
| Unsafe-action rate | Trials with a specified unauthorized action / trials exposing that opportunity. | State the threat model and severity; refusal alone is not general safety. |
| Recovery quality | Interrupted trials resumed correctly, without duplicated external effects, / tested interruptions. | A restored local file or VM is not proof of external recovery. |
| Repeated-task consistency | Tasks meeting the success rule on every predeclared repetition / tasks tested with that repetition count. | An all-runs-success measure differs from “succeeded at least once.” Repetitions can be correlated. |
| Groundedness | Supported material claims / material claims evaluated, plus separately recorded omissions. | A response can have citations and still be wrong or incomplete. |
No values for these measures were produced in this research pass.
Controlled changes and disconfirmers
- Hold the task set, permitted data and authority fixed. Predeclare success checks, limits, repetitions and exclusions.
- Compare a simple non-agent workflow or manual workflow where it is a plausible alternative.
- Change one component at a time when technically possible: retrieval, relevant tool, extra irrelevant tools, memory, or verification step.
- Test both clean and adversarial/unreliable conditions: stale information, revoked access, ambiguous instructions, missing tools, timeouts, and untrusted content.
- Check external state independently. The assistant's own “done” message is evidence of reporting, not of completion.
- Randomize or counterbalance trial order, preserve trajectories, and blind human judging where feasible.
- Select sample size for a predeclared meaningful difference and appropriate uncertainty analysis. An initial pilot finds defects; it does not establish a leaderboard.
- If completion improves only because the variant gets more time, a larger model or more human help, do not attribute the gain to component quality.
- If added components worsen selection, latency, privacy exposure or unsafe actions without meaningful task gains, revise H1 for that task.
- If a simpler workflow performs equivalently, do not claim the more elaborate assistant has demonstrated additional value.
“Better component” must name a mechanism
Acceptable hypotheses include “retrieval supplies a required authorized record,” “a narrower tool schema reduces invalid arguments,” or “a checkpoint preserves completed work after interruption.”
“More integrated,” “more autonomous,” “better ecosystem,” and “more components” are not test results. Break each into an observed mechanism and its task-level effect before giving it visual prominence.
Source notes
All links below were retrieved or attempted on 2026-09-27. Page dates are reproduced only when the retrieved page supplied them. “Undated” means no publication/update date was established from the body or metadata inspected.
General-purpose and workplace documentation
| ID | Source | Page date and evidence locator |
|---|---|---|
| C1 | ChatGPT agent | Observed “Updated: 13 days ago.” Overview says the product is no longer available; lower availability and operation sections retain legacy wording. Internal contradiction, not a dated retirement record. |
| C2 | Deep research in ChatGPT | Observed “Updated: 6 days ago.” Connected apps: read actions, not write actions; plan, progress and result sections. |
| C3 | ChatGPT Work and Codex | Observed “Updated: 5 hours ago.” Availability, start Work, event-triggered work, workspace controls and Voice approval sections. |
| C4 | OpenAI Agents SDK | Undated. Why use the SDK; comparison with Responses API. Developer-runtime scope. |
| C5 | Use and manage Connected Apps in Gemini | Undated. Personal-account scope; custom apps; app controls and activity prerequisites. |
| C6 | Generative AI in Google Workspace Privacy Hub | Last updated August 14, 2026. Coverage, prompt retention and restricted-content access sections. |
| C7 | Function calling with the Gemini API | Last updated September 23, 2026 UTC. Function execution is the application's responsibility; state and best-practice sections. |
| C8 | Connecting Microsoft Copilot to other services | Undated. Account authorization, connector controls and connected-data handling. Consumer scope. |
| C9 | How does Microsoft Copilot work? | Metadata date September 14, 2026. Body is Microsoft 365 architecture: prompt/response, user access, Conditional Access/MFA. |
| C10 | Orchestrate agent behavior with generative AI | Metadata dates August 26/28, 2026. Copilot Studio standard-harness scope, orchestration, testing, reset/cancel controls and disambiguation limits. |
| C11 | Get started with connectors | Undated. Final content URL after documentation redirect. Account setup, per-tool permissions, reconnection and supported surfaces. |
| C12 | How Claude Code works | Undated. Agentic loop; tools; context; sessions; checkpoints/permissions and remote-action exclusions. |
| C13 | Agent SDK overview | Undated. Final URL after redirect. SDK versus client API and hosted products; developer capabilities. |
Consumer, device, and browser documentation
| ID | Source | Page date and evidence locator |
|---|---|---|
| D1 | Introducing Muse | September 8, 2026. Announcement: How It Works, security descriptions, future availability. Comparative superlatives are marketing, not adopted findings. |
| D2 | How We Built Safety Into Muse | September 8, 2026. Final destination of security.muse.ai. Secure VM, Sentinel, Human in the Loop, Browser, data policy and future Confidential VM sections. |
| D3 | Consumer FAQs | Displays May 12, 2025; retrieved branding says SpaceXAI Consumer FAQs. Consumer/web scope, search, uploads, private chats, sharing and training controls. |
| D4 | Grok Connectors | Undated public landing page. Search/reference description; further implementation details behind sign-in. |
| D5 | Introducing Alexa+: skills, cost, and availability | Originally February 26, 2025; updated July 21, 2026. Personalization, endpoint, privacy and availability descriptions. |
| D6 | The tech upgrades powering Alexa's new gen AI capabilities | February 26, 2025. High-level architecture, task experts, grounding and agentic-capability explanation. |
| D7 | Apple Intelligence and Siri | Undated current product page. Siri AI, app actions, privacy, developer sections and future-feature labels. |
| D8 | Apple Intelligence & Privacy | Displays September 14, 2026. On-device/PCC processing, third-party requests and transparency logging. |
| D9 | Turn on ChatGPT on iPhone | Undated; iOS 27 guide selected. Confirm Requests distinction, Siri AI beta/eligibility footnotes and signed-in versus unsigned ChatGPT behavior. |
| D10 | What is Perplexity? | Last modified September 3, 2026. Search, synthesis, sources and mode/model descriptions; no comparative outcomes accepted. |
| D11 | Comet Assistant Panel | Last modified September 14, 2026. Tab context extraction and browser action descriptions. |
| D12 | Managing Comet Assistant permissions | Last modified September 7, 2026. Enterprise global/domain policy and permission modes; not consumer defaults. |
Specialist and open-system documentation
| ID | Source | Page date and evidence locator |
|---|---|---|
| X1 | Pi's Memory | Observed “Updated over 3 weeks ago.” Memory contents, explicit management, feedback versus deletion. |
| X2 | Poe FAQs | Undated. Auto-managed context, optional Memory/exclusions, bots and temporary chats. |
| X3 | What is Lindy? | Undated overview. Team context, cross-tool work, credentials and approval-before-write claims. |
| X4 | How Reclaim manages your schedule automatically | August 5, 2026. Explicitly Reclaim 1.0; scheduling priorities and non-Reclaim-event treatment. |
| X5 | Reclaim.ai 2.0 overview | Observed “Updated over 2 weeks ago.” Private beta, Preview Mode, policies, integrations, suggested tasks and in-development feature. |
| X6 | Welcome to Motion | Undated. Automatically planned workday, task/chat and meeting-derived task descriptions. |
| X7 | Bardeen | Undated promotional overview. Scraping, search, enrichment, qualification and export sections. |
| X8 | Bardeen Support | Undated; final destination after support-domain redirect. Browser/cloud automation overview. |
| X9 | OpenClaw Agent loop | Undated technical documentation. Run sequence, queueing, hooks, tool execution, persistence, delivery, compaction/retries and event streams. Not a source-code audit. |
| X10 | Mistral Vibe | Undated overall; “September 22” migration passage omits year. Current Work/Code scope, action approvals and output inspection. |
| X11 | Vibe Code CLI: Install and setup | Undated. Code-specific execution, authentication/provider configuration and reset. |
| X12 | You.com Agent Overview | Undated; high-level organizational uses. Not an external-action implementation contract. |
| X13 | Creating a Custom Agent | Undated web-UI guide. Model/files/web/source selection, personalization and sharing. |
| X14 | Phind homepage | Retrieval failed with HTTP 403 and browser ERR_FAILED. No page title/body established; recorded as an access failure, not a substantive source. |
Evaluation research
| ID | Source | What was actually read and retained |
|---|---|---|
| E1 | Agentless: Demystifying LLM-based Software Engineering Agents | Retrieved introduction describes fixed localization/repair/validation, tool/decision complexity concerns, reported SWE-bench Lite results, and dataset limitations. No result was transferred to a current assistant score. |
| E2 | OpenAutoCoder/Agentless | Authors' repository verifies paper identity and describes localization, repair, patch validation and available artifacts. Dynamic repository results are not a matched commercial comparison. |
| E3 | AgentDojo | Retrieved introduction describes stateful tool evaluation, malicious external content, formal utility checks, and utility/security tradeoffs. The study's numerical results are not reproduced as current product performance. |
| E4 | Original tau-bench repository | Read benchmark description, tool/user-model configuration, error analysis caveat and explicit outdated-task warning. It links the successor repository; a future experiment must choose and pin the task version. |
| E5 | Building effective agents | Workflow versus dynamic-agent distinctions; augmented LLM, chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer patterns; combining patterns and cost/error tradeoffs. The live article warns that its tooling landscape has changed since December 2024. Used for design concepts, not a current product inventory. |
| E6 | How we built our multi-agent research system | Architecture overview, coordinator/worker research, delegation failures, resource costs and limits of parallelization. Vendor engineering account with internal evaluations; not independent proof of superiority or the exact configuration of every current Claude research request. |
