Skip to content
Harkster

AI DATA CONSULTING FOR INSTITUTIONAL INVESTMENT FIRMS

Your data, in AI-ready shape.

Harkster helps institutional investment firms convert their macro research and data sources (email, files, APIs) into clean, embedded, indexed data ready for downstream AI. We deliver the data foundation; your teams build the proprietary last mile.

Built on the Harkster Kernel. Deployed in your environment. Your data never leaves your infrastructure.

THE PROBLEM

Everyone wants AI on the desk. Nobody's data is ready for it.

Institutional investment firms sit on an enormous, valuable stream of proprietary research and market intelligence, but it arrives as unstructured email, PDFs, and portal downloads scattered across inboxes, chatrooms and folders.

Our analysis found some portfolio managers try to process more than 500 research items per day, yet on a quiet market day day only 20–30% of research is ever read. Trade ideas sit unseen in unread content. Hours are lost searching chats, inboxes and portals.

Meanwhile, retail AI tools have been taken off many trading desks: compliance can't allow research content to be pasted into consumer LLMs that may train on it. The appetite for AI exists, but the data is not in a usable or governable form, and that is the real blocker.

500+

research items a portfolio manager tries to process per day

70–80%

of research is never read, even on a good day

Unstructured & scattered

Research arrives as email bodies, PDF attachments, charts and tables, locked in formats AI can't reliably consume, spread across inboxes, drives and portals.

Ungoverned AI usage

Market participants copy and paste research into retail LLMs, creating IP-leakage and compliance risk. Firms need a controlled path to AI, not a ban.

Stalled AI initiatives

In-house AI projects burn months on data plumbing (parsing, cleaning, deduplication, indexing) before a single workflow ships. Most never get past the plumbing.

The last mile of AI is proprietary. The first mile shouldn't have to be.

THE SOLUTION

We get your data AI-ready. You build the edge.

Harkster is a consulting practice built around a proven piece of technology: the Harkster Kernel, a data adapter for institutional research that was developed and battle-tested on live hedge fund trading desks.

We work with your team to connect your data sources, deploy the Kernel inside your environment, and deliver a continuously updated stream of cleaned, embedded, indexed, AI-ready data, plus derived intelligence layers built on top of it.

From there, your firm owns the last mile: your engineers and investment teams define and manage in-house workflows that match your bespoke, proprietary processes, without exposing your methods to anyone, including us.

Harkster delivers

Email

Files

APIs

Harkster Kernel

Clean · Embed · Index · Derive

You own

Your Workflows

In-house, proprietary, yours

01

Connect

We wire in your sources: Outlook email via Microsoft Graph, file drops, and third-party or internal APIs.

02

Convert

The Kernel screens, parses, extracts, cleans, normalises, deduplicates, embeds and indexes everything, turning raw content into structured, queryable, AI-ready data.

03

Compound

Beyond raw data, the Kernel produces first- and second-derivative intelligence: topic summaries, trade ideas, watchlist sentiment, trending themes, idea consensus, ready for your downstream use.

THE ENGINE

The Harkster Kernel: a data adapter for institutional research.

The Kernel is the core of every Harkster engagement. It ingests your existing research inputs (bring your own research) and outputs three layers of data, each ready for downstream consumption via API, MCP, or direct integration into your systems.

Raw, first-derivative, second-derivative

Layer 0 · AI-Ready Raw Data

Every source item, cleaned, embedded and indexed. Structured JSON with full provenance and citations. This is the foundation: query it, retrieve against it, feed it to any model or agent you run in-house.

Layer 1 · First-Derivative Intelligence

Extracted signal from each individual item:

Topic summaries

Every item summarised and decomposed into its core topics and claims.

Trade ideas

Explicit and implied ideas extracted, classified by asset class, direction and conviction, with supporting quotes.

Watchlist items

Mentions of the assets and entities you track, each scored for sentiment.

Event extraction

Items mapped to the risk events they relate to.

+ More

More extraction types available, and new extractions can be built around your specific content.

Layer 2 · Second-Derivative Intelligence

Signal aggregated across items and time:

Trending themes

Derived from topic summaries: what's dominating your research feed over the past 24 hours, with narrative summaries and source evidence.

Trade idea consensus

Derived from extracted trade ideas: normalised and aggregated views by asset, showing breadth, split and net direction.

Watchlist sentiment EMA

Derived from watchlist items: an exponential moving average of sentiment per asset, turning scattered mentions into a trackable signal.

Event Preview / Review

Derived from event extraction: synthesised pre- and post-event reports bringing together consensus views, scenario analysis, catalysts and positioning around a specific risk event.

+ More

More aggregate views available, and bespoke views are built around what your desk tracks.

All three layers stream into your systems via the Harkster API / MCP, clean, classified, cited, and ready for whatever you build next.

HOW WE WORK

From discovery to production: a clear path to AI-ready data.

Every engagement pairs Harkster's financial-markets and engineering expertise with your teams. We've spent years building this pipeline for the most demanding research consumers in the market; you get that experience applied directly to your data landscape.

1

Discovery & Assessment

We assess your current workflows, data sources and systems landscape to uncover practical AI use cases. Together we prioritise opportunities by business value, technical feasibility and speed to deployment.

2

Kernel Deployment

We deploy the Harkster Kernel into your enterprise cloud environment and connect your sources: email via MS Graph connector, files, and APIs. Hands-on support across infrastructure, security, integration and connectivity.

3

Data Validation & Tuning

We tune extraction, tagging and enrichment to your content and your language, applying deep financial-markets experience so outputs fit your users, instruments and operating environment.

4

Enablement & Last-Mile Handover

Your teams take the AI-ready data and build. We provide onboarding, training and documentation so your engineers can develop and manage proprietary workflows in-house, keeping your process IP yours. Optionally, we build example assistants with you (see use cases below).

The principle: we industrialise the first mile so you can keep the last mile private.

Clients create and run bespoke last-mile workflows without exposing proprietary methods, preserving the information edge for your investment teams.

BUILT ON THE KERNEL

See what AI-ready data becomes on the desk.

The Kernel is the foundation: these are the kinds of assistants investment teams run on top of it. Some we've built with clients through our secondary consulting offering; all consume the same three layers of Kernel data. Use them as a starting point, or as proof of what your own last mile could look like.

Event Risk Calendar

"Consensus in a click."

How long does it take you to prepare for a central bank meeting? This assistant pulls together all research relating to a specific risk event in one click. Pre-populated with a client-defined calendar, select an event to see the most relevant research, extracted topics, and a concise summary with citations. Everything needed to price the risk, in one place.

Layer 0 raw data + Layer 1 event extraction and topic summaries.

Trending Themes

See what's driving the conversation.

Automatically surface the key themes emerging across a research feed in the last 24 hours. Related insights are grouped, the narrative summarised, core claims extracted, and every theme points back to the source evidence behind it. PMs get the shape of the day's flow in minutes, not hours.

Layer 2 trending themes, derived from Layer 1 topic summaries.

Trade Idea Consensus

"Quickly scan for clear, consensus views."

No more skim-reading every piece looking for ideas. Trade ideas are extracted from every item, then normalised and aggregated into a consensus view, by asset, with breadth, split and net direction, populated only from the client's own content.

Layer 2 trade idea consensus, derived from Layer 1 trade idea extraction.

Q&A

Ask questions of your research feed.

The research reservoir becomes a database to mine rather than something to tame with rules and folders. Users ask questions in plain language and get cited answers drawn exclusively from their own research: insights and perspectives that would otherwise stay buried.

Layer 0 raw indexed data via the Kernel's hybrid retrieval, reranking and citation layer.

Concepts

What will you build?

The assistants above are just some examples of what has been built so far. With AI-ready data streaming into your systems, the design space is open: recap agents for time off desk, scheduled proprietary tasks, sentiment overlays for systematic inputs, compliance screening, cross-domain search.

Talk to us about your use case

Book a Consultation

TRUST

Custom, secure and private, by design.

Your Cloud Environment

Harkster Kernel

Deployed inside your infrastructure; data never leaves.

Stateless LLMs

No prompts or generations stored in the model.

AES-256 at rest · 256-bit TLS in transit · no training on client data

Harkster: consulting, no data access

Deployed in your environment

The Kernel runs "on-prem" within your own cloud infrastructure. Your data never leaves your environment.

No training on your data

Client data, prompts and generations are never used to train, retrain or improve base models, and no publisher IP is used either.

Stateless models

The LLMs used are stateless: no prompts or generations are stored in the model.

End-to-end encryption

Data encrypted at rest with AES-256 and in transit with 256-bit TLS.

Governed access to AI

Managers can screen the content accessing the LLM. Each client operates in their own fully independent environment.

Enterprise-ready integration

Connect securely with Microsoft Graph and other enterprise systems using access controls configured around your organisation’s policies.

WHO IT'S FOR

Built with, and for, institutional investment teams.

Harkster's technology was developed on live macro trading desks and refined with portfolio managers, traders and technologists at institutional investment firms. We speak both languages: the CTO's need for clean architecture, governance and integration, and the desk's need for speed, signal and edge.

For technology leaders

A proven, deployable data adapter instead of a multi-quarter, million-dollar internal build. Corporate email connectivity, clean APIs, full governance, and a partner with deep experience in both financial markets and AI engineering.

For investment teams

Research that finally works for you: themes, ideas, sentiment and answers surfaced from your own content, feeding assistants and workflows shaped around how you actually trade.

"Harkster is my go-to each morning for fast information extraction from 100s of emails."

— Trader, Foreign Exchange

"Event Risk Calendar is a really efficient way to summate multiple views and forecasts on data points."

— Portfolio Manager, Global Macro

"Distills a large array of views that would otherwise take considerable time to do myself."

— Portfolio Manager, Global Macro

"Harkster saves me time. Probably more than a full hour per day."

— Trader, Foreign Exchange

NEXT STEPS

Let's look at your data landscape.

Book a consultation and we'll walk through your sources, your systems and your ambitions, and map the fastest path to AI-ready data in your environment. Prefer proof first? Send us a sample piece of research and we'll return the Kernel's output within 24 hours.