Cognition · NewGlobe
A quiet interface for asking hard questions.
Conversational AI for the officials who run a national education programme, and answer to ministers for it. The screens here report on Bayelsa State, Nigeria, covering 222 schools and 41,000 pupils. NewGlobe unveiled it at the Education World Forum in London, where it drew applause from an audience of education ministers from more than 100 nations. I led the design from problem framing through to the shipped interface, working with a product director, a solution architect and the engineering team.
- Role
- Design lead
- Scope
- Research, product design, design system
- Timeline
- Nov 2025 – May 2026
- Team
- Product director, solution architect, engineers
Data, not answers
NewGlobe held one of the richest education datasets anywhere: attendance, lesson delivery and assessment captured daily across every school in the programme. The dashboards on top of it were real, and people used them.
They just answered a different question. A dashboard shows you the numbers and leaves the interpretation to you: pick the filters, cross-reference two views, work out which of the three districts moved and why. That is analyst work, and it takes time nobody in this audience has.
The people using them had a specific question, a meeting in an hour, and a phone in their hand. The gap was never access to the data. It was the distance between a chart and a sentence they could say out loud.
A minister doesn’t want a dashboard. They want an answer they can repeat in the next room.
Field note from the commercial team
Three readers, one accountability
Every design decision traced back to who was holding the phone. All three answer upward, none of them are analysts, and none of them have time to learn a tool.
They are senior people in their forties and fifties whose expertise is education policy, not software. Conversational AI is not part of how they already work, and nothing about their job has required it to be. The unfamiliarity sits with the tool, not with them.
NewGlobe held extensive persona research on this audience already, and I ran a further round of interviews on top of it. The three below came out of that work, which is why the shape of the product was settled before anything was built.
- Ministry officialsSpeak about education to cabinet, parliament and the press. Need a short, defensible answer, not a spreadsheet.Speaks in public, not in software
- Regional directorsRun schools at scale. Want their region compared to last term, before a 9am briefing.Patchy connectivity
- Programme leadsThe bridge between NewGlobe and the partner government. Brief upward, intervene downward, fast.Multi-stakeholder load
What we tried before a conversation
A conversation was not the obvious answer. Three other shapes were explored before it.
- 1
A smarter dashboard
The cheapest thing to build on what already existed, and it kept the original problem intact: the numbers are still somewhere you have to go and find. A minister in a government office does not apply filters, and should not have to.
- 2
A WhatsApp bot
It met the audience where they already were, phone in hand, between visits. It also put a government system inside a teacher’s personal WhatsApp. Neither the privacy exposure nor the question of who owned that thread had a good answer.
- 3
A desktop app built around a teacher’s persona
The closest to a product with a personality. It ran into a fear we could not design around: that anything installed and always running was NewGlobe, or the government, watching. Rather than try to mitigate a suspicion that was reasonable, we dropped the shape that provoked it.
What won was a chat at a URL. It sits behind a login on NewGlobe’s own domain, carrying NewGlobe’s identity, which answers the ownership and surveillance questions the other three shapes could not.
That login carries the access rules with it. Programme isolation is absolute: a user cannot retrieve data for a programme they are not cleared for, under any phrasing of any question, and asking sideways returns a refusal that names the programmes they do have rather than hinting at ones they do not. An official in one state never sees another state’s numbers, and never sees a control suggesting they might. Chat history is scoped to the individual, not the programme, so nobody browses a colleague’s questions.
A conversation was not a familiar form for this audience either. It is why the opening screen leads with example questions instead of an empty box: the blank prompt is the part of chat they cannot use.
The decision: a tool that says no
The obvious build was a capable assistant that answers whatever it is asked. I argued for the opposite, and it became the shape of the whole product: Cognition answers about programme data and declines everything else, out loud, citing its scope.
Refusal is usually treated as a failure state to be minimised. Here it was the feature. An assistant that visibly declines is one an official can put in front of a minister without rehearsing it first.
- 1
Trust
Officials learn the edges of the tool in one turn. No invented policy, no opinion on a rival programme, no drift into territory nobody signed off.
- 2
Procurement
A narrowly scoped, predictable system is far easier to put through a government IT review than a general-purpose assistant.
- 3
Composure
It refuses in plain language, with no apology and no hedging. An official reading the reply aloud in a meeting is not embarrassed by it.
When it does not know
Declining a question it was never for is the easy half. The harder half is a fair question it cannot answer well, because that is the answer an official repeats to a minister without knowing anything went wrong.
The rule underneath the product is closed-loop grounding. Cognition does not search the open internet, speculate or fabricate. Every answer is generated from authorised programme data, and where there is no data to answer with, it says so plainly rather than guess. Internet access was ruled out of the release for exactly this reason: an ungrounded answer is worse than no answer to someone who has to repeat it in a meeting.
- 1
No data for the question
It says what it checked and what it did not find, and it does not present an empty result as a meaningful zero. A period with no data is not a term where nothing happened.
- 2
A question it cannot answer
A plain explanation of why, and a suggestion for rephrasing or narrowing. No fabrication, and no half-answer that leaves the reader to work out which part to trust.
- 3
Data outside the user’s programme
It declines and names the programmes the user does have. It never returns partial data, and never implies a figure exists but is being withheld.
- 4
Something goes wrong on our side
It says so and offers a retry on the same question. Errors do not arrive dressed as answers.
The sharpest version of this was a decision to carry less. Assessment and teacher observation data existed, but it had not been prepared in a form Cognition could answer accurately, and assessment is the highest-stakes thing this audience asks about. Both were switched off before launch on the grounds that they were doing more harm than good. Shipping a narrower product was cheaper than shipping a confident wrong answer about a child’s learning.
What a good answer looks like
Answers are set as prose at a comfortable measure. A chart where the question is about a trend, a table where the content is genuinely tabular, and a sentence after either one saying what it means. No chat bubbles, no gradients, no chrome signalling “AI”. The register borrows from a well-set annual report, because that is the document this audience already trusts.
The voice was mine to set. I read what good practice actually looked like for assistants of this kind, and wrote the guidelines the answers are held to: what it calls things, how much it hedges, how it declines, and how long an answer runs before it stops being useful.
What it calls things is not one decision, though, because the same words do not mean the same thing across programmes. Streams in one place are arms in another. Pupils in one are students in another. Head teacher, head master and school leader are three names for a single role, and last term resolves against a different calendar in each state. Cognition normalises what you type to a canonical term before it goes anywhere near the data, then answers back in the vocabulary of your own programme. None of the mapping is visible: you ask in your words and the answer arrives in them.
That mapping is doing real work. Without it, every official outside whichever programme happened to set the canonical terms would be translating their own vocabulary into somebody else’s before they could ask a question, and reading the answer back the same way.
The second decision: two modes, not three
The solution architect’s model was three modes of thinking: fast, medium and deep. It described the system accurately. It told an official nothing, because how hard a model is working is not something a user has any way to judge, or any reason to care about.
I argued for two, named for the question rather than the machine, and sat down with the product director and the solution architect until we agreed on it. The difference between them is not how long the model thinks. It is whether the answer is checked before anyone sees it. Lite returns what it generates, which suits the arithmetic most of this audience asks for: how many teachers, how many pupils, how many schools. Pro takes a further pass to verify the answer before showing it. The budgets we designed against were 30 seconds for Lite and 60 to 90 for Pro. Lite is for a question you can sanity-check yourself. Pro is for one you intend to act on.
The system picks the mode itself, and we instrumented that switch so we would find out which one people actually used. Lite was the default in practice.
Every answer carries the mode that produced it, so the reader knows which of the two they are holding. Pro shows the working: what was compared, against which baseline, in what order, and the extra step where it checks itself. Deciding how much of that to show, and when, took three attempts.
- 1
Two separate boxes
A thinking indicator above a list of steps. Accurate, and visually noisy: two containers competing before the answer had even arrived.
- 2
One merged box
Collapsed into a single panel showing the active step and a live timer. Better, but it sat open after completion and pushed the answer down the page.
- 3
Collapse on completion
It now folds itself away shortly after finishing, leaving one quiet line: how long it thought, and how many steps. The steps persist on the message, so anyone who wants the working can reopen it.
Show the working when someone asks for it. Hide it when they don’t.
Getting it used
The officials work from Android tablets. For that audience an app is an icon on the home screen, not a URL in a browser, so Cognition ships as an installable PWA rather than a native app or a plain responsive site.
It did not need inventing. NewGlobe already ran a PWA for Spotlight, the dashboard product this one answers back to, so Cognition took the same shell. The work was everything around it: 15 icon sizes, splash screens held to a 3.0-second cold start and a 1.8-second warm start, breakpoints from 360 pixels up to tablet, and a sidebar that becomes a drawer below 1024 rather than a rail.
It still needs a connection. Installed, it opens like an app; away from signal it does not work.
Voice input follows the same logic. Speech lands in the composer as editable text rather than firing off as a query. Speech recognition is trained overwhelmingly on accents that are not these officials’, and a mis-heard question returns a confident answer about a district nobody asked about. Showing the transcript first puts the correction in the user’s hands instead of making them argue with the machine.
The tool had to be in officials’ hands quickly, and an in-product tour was the wrong instrument twice over: it would have cost time we did not have, and this was not an audience likely to follow one. So we taught it outside the product. Marketing ran live online training sessions for the officials on what Cognition is, what it is not, and what to ask it. I made the material they taught from.
That framing came straight out of the scope decision. A tool defined as much by what it declines as by what it answers is a tool you can teach in a single session, because the boundary is the lesson.
Install prompt
Installed
Sidebar as a drawer
Where the research ran out
Two things about this audience were settled before any of it was built. The persona research inside NewGlobe, and the interviews I ran on top of it, said these officials would not meet a blank chat box halfway. So the opening screen was never a blank box. I proposed leading with example questions, and the team agreed there should be handholding rather than an empty screen with no direction. That call held.
What the research did not catch was smaller and more ordinary. An early version put logout behind a click, the way most current interfaces do. Watching officials use it, nobody found it: this is an audience whose habits were formed when logout was always on screen. It is permanently visible now.
The same thing turned out to be true of two more controls in that corner of the screen. The theme switch was a sun and a moon, which is the convention everywhere and meant nothing here, so it became the words Light and Dark under a heading that says Theme. The control that collapses the sidebar was doing the same quiet disappearing act, and was made explicit alongside them.
The parts I had reasoned about hardest held. What broke was the conventional furniture around them.
Before: a moon, and no way out
After: both spelled out
How accurate it had to be
The floor was 80% accuracy measured against Spotlight, the dashboard product this audience already treats as the record, with 85% as the number we were aiming at. The test was not whether an answer read well. It was whether it matched the number the organisation had already agreed on.
It was calibrated before any official saw it. The first release went to about 25 people inside NewGlobe. 10 came from the executive and commercial side, asking the questions they would really ask. 15 came from the technology and Cognition teams, whose job was to check answers against Spotlight and file the ones that did not match. Staff attendance took the longest to come up to standard. Officials followed in a staged rollout, roughly 20 per programme across Bayelsa and Jigawa states in Nigeria, and Liberia.
80% also means 1 answer in 5 is wrong, and the interface has to carry that honestly. So every answer shows which mode produced it. Every answer can be marked useful or not. Pro exists for the questions where being right matters more than being quick. Nothing on screen claims more confidence than the system has earned.
The feedback control was the whole point of that first release rather than a courtesy. Without a way to flag a wrong answer inside the product, accuracy problems arrive as screenshots in chat threads, stripped of the one thing that makes them fixable: which question, in which mode, in which session.
Those 25 people were also how the design got tested. There is no clean usability protocol for an interface wrapped around an answer that differs every run, so we ran it as a weekly loop instead. The group met each week, what they hit became tickets, and the fixes went into the next build. My job was to turn what came out of that hour into design fast enough to be in it. Underneath that ran a daily conversation with the solution architect and the product director, where we decided what went out next and what waited.
Most of what the product became came out of those weeks rather than out of a document. Three of the larger ones:
- 1
Search, and how history is grouped
Whether past conversations should be grouped by day, week or month, and what search across them needed to match, were both settled by watching people fail to find something they had asked before.
- 2
The programme selector, removed
I had designed Cognition in every programme’s brand colours, NewGlobe’s included, and put a selector in a modal at the start. The loop killed it. The programme belongs in the URL, and each person is sent to the one they are assigned to, so most users never see a control at all.
- 3
Dark mode, argued down
The proposal was a dark theme built from NewGlobe’s brand colours. I argued it would fail on contrast before it reached anyone, and we shipped a conventional dark theme instead. Every screen was redrawn for it.
The same weeks produced the smaller things that make it usable: the feedback control itself, and tooltips on the controls for people meeting these options for the first time. Starring and renaming a conversation came out of the loop too, designed and not built before I left.
What shipped, and what did not
Cognition was unveiled at the Education World Forum in London in May 2026, the largest annual gathering of education ministers in the world, as part of NewGlobe’s enterprise AI suite. Their announcement page carries the film of the unveiling.
At the point it was shown, the interface, the streaming, the sessions and the design system were real and stable, with the data layer held behind a single swap-in point. What follows is what existed then.
Built
- Conversational surface with token-by-token streaming
- Persistent sessions with search and auto-titling
- Lite and Pro modes with reasoning disclosure
- Scope guard: out-of-scope questions decline and redirect
- Voice input for use between school visits
- Self-generating follow-up suggestions
- Programme isolation: a user cannot reach data they are not cleared for, under any phrasing
- Light and dark themes on a shared token set, WCAG AA contrast in both
- Published design system with an accessibility gate: keyboard reachable, screen-reader labelled, never colour alone
Scoped, and next
- Surfacing low confidence on a single answer, beyond the mode label
- Assessment and teacher observation data, switched off rather than answered badly
- Offline. It installs like an app but still needs a connection
- Citations on the face of an answer. Designed, not shipped
- Local languages. The interface is English, and speech recognition is the weaker half of that
- Starring and renaming a conversation, so a piece of analysis can be found again
Every figure in these screens is synthetic, and the screens are the design files rather than a running build. The partner data behind them is not mine to publish.