AI · the glass box

Wired into every page.Cited every time.

AI is the backbone of the terminal – reading every page, calling every tool, and binding filings, quotes, flow, macro, and your portfolio into one answer, with every claim linked back to a source.

  1. 8 AI surfaces shipped, one more in the lab, each native to the page it lives on: alert triage, classification, Morning Brief, ask the market, deep research.the suite →
  2. 99.2% citation accuracy across the eval set; every claim in an answer carries its source chip.the evals →
  3. 4 providers, one interface · ~1.8s median end-to-end, routed to the smallest sufficient model.the routing →

One answer, traced

One question about NVDA, answered in the workspace

An illustrative assistant answer in the terminal workspace: one question about NVDA's last earnings, answered in prose with a source chip after every claim. Selecting a citation highlights its source in the ledger beside the answer. Answers are generated at query time in the product, not hard-coded.

workspace / NVDA · assistantroute claude-sonnet-51.8s
What changed at NVDA in the last earnings?

Data-center revenue came in +94% YoY and Hopper sell-through is tightening into the back half . Management softened its China guidance while leaving FY27 capex plans unchanged .

On the guide: gross margin holds near 74% through FY27, networking crossed a$10B quarterly run rate , and a new $50B buyback authorization lands as a floor under the print.

Since the print: a four-officer insider cluster, $312M in sells, and net inventory days +6 QoQ.

grounded · every claim carries its source5 sources · 9 citations · 1.8s
Fig. 1 · one answer, traced · illustrative composite; in the product, answers generate at query time from live rows

The suite · who runs what

Every surface, organized by the model that runs it.

Not a chatbot bolted on the side. Each surface is native to the page it lives on, and each is routed to the smallest sufficient model. Here is who runs what:

gpt-5.6-luna
~$0.0008/queryavg 0.6s

Fast classification and single-table lookups, where latency matters more than synthesis. OpenAI's luna runs point; Sonnet 5 is the cross-vendor fallback, so one vendor's outage cannot take the tier down.

  • Smart alerts live

    Classification filter sits between raw events and your inbox: price, filings, insider, options unusual flow. Only the relevant ones fire.

  • Zero-shot classification live

    Filings, news, insider transactions, 8-K events labeled with confidence on ingestion so every downstream surface can filter by meaning, not keywords.

Sonnet 5
~$0.012/queryavg 1.4s

Cross-table synthesis with citations. The default routing for Pro members; carries most of the conversational surfaces.

  • Morning Brief live

    Pre-open synthesis of overnight moves, scheduled earnings, macro prints, and portfolio-relevant filings: one scannable brief before 9:30 ET.

  • Ask the market live

    Conversational Q&A with streaming responses. Tool-calling pulls quotes, filings, options flow, and news in-line while the answer is being written.

  • Stock intelligence live

    Per-ticker synthesis on /stock/:symbol. Why the stock moved today, what the filings say, what the options are pricing, answered before you scroll.

  • Living Analyst Notes live

    Standing AI coverage on the top 100 US large-caps (thesis, bull case, bear case, catalysts, risks) kept fresh as filings land. Layer your own thesis on top and the system flags evidence.

Opus 4.8
~$0.10/queryavg 4.2s

Multi-source thesis-level reasoning. Reserved for long-form work that justifies the cost and latency.

  • Deep research beta

    Long-form reports on a ticker, theme, or question. Multi-source synthesis, cited throughout, written for a human who has an hour to read instead of a minute.

  • Agentic research soon

    Multi-step, tool-calling agent that plans the research, fetches what it needs, and hands you a reasoned answer. In the lab, not shipped.

Voyage
embedding onlyindexed nightly

Embeddings, not generation. Powers the semantic-search surface that finds prose the keyword search misses.

  • Semantic search beta

    Find the 10-K paragraph you half-remember. Voyage finance-tuned embeddings index filings and news into a shared space; queries return the nearest prose, not the nearest keyword.

Pipeline · behind every answer

Ask. Reason. Cite.

Every answer, including Fig. 1 above, runs the same deterministic four-phase pipeline, and you can inspect every tool call.

01

Parse

Decompose your question into structured sub-queries against the alt-data schema.

02

Retrieve

Fan out tool calls to the relevant tables; pull the smallest evidence set that answers the question.

03

Synthesize

Compose a cited answer with inline source chips. No claim without a row.

04

Cache

Result cached at the question level; identical follow-ups served instantly. Stale on data refresh.

Failure · the honest path

When the feed fails, the answer says so.

Most assistants fill a data gap with a confident guess. Convexity's contract is the opposite: a failed source returns an honest empty, the answer narrows to what it can still prove, and the sync ledger is right there in the open.

An illustrative assistant answer during a partial data outage: the insider-filings feed has failed, so the assistant reports which sources are current, states exactly when the failed feed last synced, and declines to estimate what it cannot source. A sync-ledger strip below the answer shows the failed run and the last successful one.

workspace / portfolio · assistantroute claude-sonnet-51.2s
any unusual insider activity in my holdings this week?

Options flow and congressional trades are current. On insider filings I am degraded: the Form 4 feed last synced at 13:02 ET and its 13:47 run failed upstreamsec_edgar_form4 · failed.

Through 13:02: no officer transactions in your nine holdings. Anything filed since is not in my data, so I will not guess at it. The sync retries at 14:02; ask me again after.

data_sync_log · sec_edgar_form4
13:47 ETfailedrows_added 0upstream 503
13:02 ETsuccessrows_added 37
no synthetic fallbacks · a failed feed returns an honest empty1 source degraded · answer scoped to synced rows
Fig. 2 · when the feed fails · illustrative composite; the no-fabrication rule is a code path, not a style choice

The difference · same question, different machine

A chatbot answers. A glass box shows its work.

Both produce fluent prose. Everything around the prose is the difference:

Where the answer comes from
a generic chatbotModel memory plus a web search. The trail ends at a list of links.
the glass box~40 tools over the terminal's own tables: filings, quotes, flow, macro, your portfolio. The trail is the answer (Fig. 1).
How claims are sourced
a generic chatbotCitations are optional, and the model can invent one as fluently as it invents prose.
the glass boxA source chip on every claim, each traceable to a tool-call row. Measured, and published in the evals below.
When the data is missing
a generic chatbotThe gap gets filled with the most plausible continuation, and it reads identical to the truth.
the glass boxThe answer narrows to what it can prove and says so. A failed feed returns an honest empty (Fig. 2).

Try these prompts

Start anywhere. The agent finds the rest.

Every example below is a real prompt the system handles. Numbers, citations, and tool calls are generated at query time – not hard-coded into the page.

The agent has access to ~40 tools across quotes, filings, options flow, peer comparison, Convexity Score, congressional trades, and your own portfolio. It picks the smallest evidence set that answers the question, then cites every row.

prompt historyuser@convexity
>Compare NVDA / AMD FCF yield over the last 4 quarters
table + 2 chart panels · 8 citations
>Which holdings in my portfolio have insider clusters this week?
3 hits ranked by conviction · 12 Form-4 citations
>Summarize the 8-K language shifts in DHR's last filing
diff vs prior 8-K, highlighted · 2 citations
>Does the recent congressional NVDA call-options cluster suggest tail risk?
4 PTRs surfaced, weighted by role · cited
>How does NVDA's Convexity Score compare to its semis peer set?
peer cohort + percentile + sparkline · methodology linked
>Run a deep-research report on AVGO supply-chain exposure
running Opus + 14 tool calls so far…

Model routing · cost

The right model for the job.

Tier-aware routing: most queries served by the smallest sufficient model, with automatic upgrade when the agent flags low confidence. The numbers below are real internal cost metrics, not vendor list prices.

classificationgpt-5.6-lunaopenaifallbacksonnet-5anthropic
simple lookupgpt-5.6-lunaopenaifallbacksonnet-5anthropic
conversationalsonnet-5anthropicfallbackgpt-5.6-lunaopenai
synthesisgpt-5.6-lunaopenaifallbacksonnet-5anthropic
deep researchopus-4-8anthropicgroundingsonar-properplexity
semantic indexvoyage-finance-2voyageembeddings · nightly
TierTop modelLatency (avg)Cost per queryMonthly queriesDeep research
Free $0gpt-5.6-luna0.6 s$0.0008100
Pro default $29 / mosonnet-51.4 s$0.0121500
Pro Plus $79 / moopus-4-84.2 s$0.101,00020 / mo

Every call is recorded by surface, model, and token count in theAI usage ledger. Per-user monthly caps are enforced server-side; the ledger is what backs the cost-transparency claim.

Four providers, one interface: OpenAI's gpt-5.6-luna runs the classification and synthesis lanes, Anthropic's Claude carries conversation and deep research,Voyage embeds the filings corpus, and Perplexity grounds web search inside answers. Every lane keeps a cross-vendor fallback warm.

Prefer code? The same tool layer is public over MCP on the Developer and Quant tiers. Queries reach providers as anonymized context, never PII (privacy policy).

Evals · current snapshot

The numbers behind the claims.

Eval metrics pulled from the production ledger over the seven days before the snapshot date below. Citation rate is measured against an internal sampler that scores whether every numeric claim in the response ties back to a tool-call return value.

Search calls per Deep Research report

7-day median. Optimized down from 6.1 in the Phase 2 baseline by adding source-aware tool routing.

3.8target ≤ 4.5
Citation accuracy

Numeric claims tied to a tool-call return. The 0.8% failure mode is mostly rounding drift, not hallucinated data.

99.2%target ≥ 99%
Median end-to-end response

From submit to first cited token. Prompt-cache hits cut another 400ms on repeat queries inside a session.

1.8starget ≤ 2.5 s
Spend cap incidents

Per-user monthly caps enforced server-side. A user has never been surprise-charged because the ledger blocks the call before it dispatches.

0target 0

Snapshot as of ; the published numbers update when this page does, not nightly. The eval basket itself re-runs nightly, with results in the internal admin dashboard.

The boundary · drawn in code

In the box. Never in the box.

A research copilot with your portfolio attached needs a hard boundary, and it has to be drawn in code rather than copy. Both sides of it, as shipped:

in the box
  • Your holdings, as context. The portfolio is a tool the agent calls, same as quotes or filings; answers scope to what you own.
  • Every call on the ledger. Surface, model, and token count, recorded per call and priced per query.
  • Hard caps, server-side. The ledger blocks an over-cap call before it dispatches. Incidents to date: zero.
  • A source on every claim. No claim without a row; the citations are load-bearing, not decoration.
never
  • PII to a model provider. Queries leave as anonymized financial context, never personal identity (privacy policy).
  • A surprise charge. Caps are enforced before dispatch, so the bill cannot outrun the ledger.
  • An invented number. A failed feed returns an honest empty, never a plausible guess (Fig. 2 above).
  • Silent staleness. Data age is printed next to the data, on every surface, in the same visual language.

One terminal. One AI backbone. Every answer cited.

Founder pricing locked for life: $0 during the waitlist, $29/mo when we open.