Data Science & Analytics — London, UK

Feruza Kachkinbayeva

Messy data.
Governed pipelines.
AI that cites its sources.

scroll

Ask me anything

Grounded in real work. Speaks in my voice if you want.

feruza-agent@london:~
◆
$

Location Intelligence: Cafe Site Selection

SLA Master's Award 2024

4,835 LSOAs. AHP-weighted opportunity scores. Built for my MSc thesis.

loading map_

AHP weighting synthesises 12 site-suitability factors (footfall potential, competition density, transport access, and demographics) into a single opportunity score per LSOA. Consistency ratio: 0.069.

University of Greenwich · MSc Thesis 2024
Distinction

Location Intelligence — Café Site Selection

SLA Masters Award 2024 · 2nd Place

Problem

74% of new cafés in the UK fail within five years. Site selection costs £50,000+ to get wrong and most decisions rely on gut instinct and borough-level demographics that miss how dramatically conditions vary within short distances across London.

What it does

A three-task data-driven framework across 4,835 London LSOAs at granular sub-borough level. Task one: predict café success potential by area. Task two: predict commercial rent prices. Task three: find the intersection where success potential is high and rent is below market rate — that is where you open.

Key finding

Public transport accessibility dominated the model (16.92% AHP weight), above median house price, demographics, or foot traffic proxies. The more commercially valuable output was identifying emerging neighbourhoods with medium-high success scores not visible in traditional market research — areas that conventional consultancy would miss entirely.

Scale

16,361-word dissertation covering multi-source data integration, AHP weighting methodology, geospatial ML modelling, and rent prediction — then synthesised into an interactive map deployed as part of this portfolio.

PythonGeoPandasscikit-learnAHPFoliumPostgreSQLGeospatial ML
University of Greenwich · Research Assistant · Jun-Aug 2024
300K+ reviews analysed

Consumer Behaviour Research · Restaurant Inclusivity

Problem

Restaurant inclusivity research usually runs on small samples or a single review platform. Dietary-need signals, allergy, religious, lifestyle, sit scattered across Google Maps, Facebook and TripAdvisor as free text, across three cities, with no shared structure.

What it does

Built a multi-source ingestion pipeline across 300,000+ reviews with automated translation, deduplication and cross-platform normalisation. Classified 57,825 dietary-preference instances into lifestyle, medical and religious segments, the same microsegmentation logic consumer-insight teams run on customer data, applied here to review text.

Key finding

Compared TextBlob against VADER sentiment scoring on the same text before trusting either one, then used N-gram extraction and LDA topic modelling to surface the dietary vocabulary underneath. Shipped as Tableau dashboards for non-technical stakeholders: the analysis only mattered once someone who had never touched the raw data could read it and decide something.

Scale

Two concurrent research-assistant positions at Greenwich's Tourism and Marketing Research Centre. A parallel qualitative study from the same period, on school nutrition, went on to win Best Paper at CHME Conference 2025.

PythonTextBlobVADERLDATableauNLP
University of Greenwich
Prototype · CIO-approved for build

HESA Stat Returns Hub

Problem

HESA statutory returns are a regulated workflow with hard deadlines. Missed sign-off gets reported to the Office for Students. The existing process ran on Banner extracts, Python and Alteryx scripts, Excel trackers, email chains. Hundreds of quality rules per cycle, no central view, no audit trail.

What it does

Governance and submission pipeline in one tool. Role-based access, invitations, append-only audit log, multi-institution dashboard with risk scoring. The pipeline handles XML upload, lxml-based XSD validation, OVT quality report ingestion, per-rule triage and team assignment, failure drill-down, and Core File generation from the 28 TSV outputs HESA returns after sign-off.

Planned

Exploring adding an LLM layer for natural-language rule queries and tolerance-request drafting, grounded in the regulatory guidance documentation.

DjangoReactTypeScriptPostgreSQLlxml
31 days · 165 commits
GitHub

LifeOS

What it is

A mobile-first personal operating system. Today view with energy tracking, calendar, goals, check-ins, analytics, and a conversational assistant with persistent memory. Full-stack, deployed, in daily use since Day 31.

Architecture decision

The memory system went through three versions in three days. Store everything. Inject everything. Score and select. The first two were thorough. Only the third was useful. The problem was never storage — it was knowing what matters right now. The assistant uses selective injection: memories are scored against the current context and only passed to the model above a relevance threshold.

What I learned

Working and right are different things. The lesson wasn't to slow down. It was to know what I was optimising for before I started.

FastAPIReact NativePostgreSQLOpenAIRailway

What I'm working through

A running log of what I'm building, what broke, and what I changed my mind about.

thinking.log0 entries
loading_

Working with me

How I think

Most of what I build starts messy: a regulator's guidance scattered across documents, six datasets that don't share a key, a day with no shape until I give it one. I'd rather spend the first effort on structure than output.

What I'm building toward

I've had a pipeline run clean for a cycle then break because a source changed shape, and an AI agent whose memory got worse the more I gave it, until I made it selective. Getting something to work once is easy; the real problem is making it hold up on the tenth run.

What kind of work I want

Work where the output changes a decision, not a pipeline that just runs quietly. I turned a two-week manual return into two hours, and the business noticed the saved time, not the code behind it. I want more of that: analysis a non-technical room can act on, not just run.