Orbira Labs Join the waitlist

Category: Engineering

  • Designing an assistant that explains itself

    Designing an assistant that explains itself

    A tool that acts for you has to answer one question at any moment: why did you do that? If the answer is “the model decided”, the tool is not ready.

    Explanations are a design constraint, not a feature

    We treat explainability as a constraint on the whole system. Every step in a workflow has to declare its inputs, its action and its output. Every run stores them. This costs us flexibility. Some things that would be easy to do invisibly are hard to do visibly. We accept that cost.

    What gets recorded

    For each run we keep:

    • What started it, such as a schedule or an event.
    • The steps in order, with the time each began and ended.
    • What each step read, summarised and linked back to the source.
    • What each step wrote, and where.
    • Which approvals were requested, and who answered.

    Why summaries link back to sources

    When a generated summary says “three tasks closed in the billing project”, you should be able to click through to those three tasks. Otherwise you are being asked to trust prose. Linking keeps the summary honest and makes errors visible.

    Mistakes are the real test

    An explanation system proves its value when something goes wrong. A good one lets you find the failing step in under a minute, see what it was given, and fix the instruction. A weak one leaves you guessing.

    Trade-offs

    Detailed logs mean more data stored. We keep logs to what is needed to explain a run, and you can delete a run history at any time. We cover how we handle data on the privacy page.

    What we want next

    We would like a plain-language “why” on every step, written for a non-technical reader. That is harder than it sounds, because a short explanation is easy to get subtly wrong. We would rather ship a precise one that is a little dry.

  • What a good run log looks like

    What a good run log looks like

    Logs are usually written for engineers. A run log for an assistant has a different reader: someone who wants to check the work, quickly, without learning the internals.

    The reader’s questions

    1. Did it run when it should have?
    2. What did it do?
    3. Did it ask me for anything?
    4. Did anything fail?

    A good log answers these four in the first screen.

    A sketch

    Friday status summary, run 14 · 15:00 · succeeded

    • Trigger: schedule, Fridays at 15:00.
    • Read 12 closed tasks from the project tracker.
    • Grouped them into 3 projects.
    • Drafted summary (5 bullets).
    • Approval requested at 15:01, approved at 15:09 with one edit.
    • Posted to the team channel at 15:09.

    This is an example of the format, not real data.

    What belongs

    Counts, names of sources, times, approvals, and links to the underlying items.

    What does not belong

    Raw model output in the main view, internal identifiers, and anything that makes the reader scroll to find the result. It can live behind a “details” link for the person who wants it.

    Failure should be loud and specific

    When a step fails, the log should say which step, what it was trying to do, and what the likely cause is, for example that a connection lost its permission. Then it should offer the obvious next action, such as reconnecting.

    Retention

    Logs are useful, but they hold information about your work. We let you choose how long run history is kept and delete it whenever you like.