Accepted at IEEE OCEANS 2026

Nine minds.
One system that
understands animals.

An octopus has nine brains — one central, eight in its arms. We build the models that read what all of them are doing, straight from the cameras a facility already runs.

tank cam 09 · live on-device
octo-mask-lraspp · 3.2 M params · no prompt at inference · mask area 6.2 % of frame
Peer-reviewed at IEEE OCEANS 2026Monterey · accepted IEEE OCEANS 2025DISC-O closed loop IEEE ISMAR 2025Cephalopod AR TU Grazresearch partner
892 h
of continuous footage analysed end to end
5
open models, from 0.6 MB to 1.7 GB
3.2 M
parameters in the deployed vision model
$0.0006
to label a clip, five passes, no GPU on site
Our models

Five models. One animal's
entire behavioural record.

A frontier vision-language model labels the footage once, offline. We distil it into small models that run continuously on commodity hardware — and we open them.

Aquarium camera frame in which the octopus is climbing the tank, with the model's segmentation mask painted over its body and every arm in teal. mask 6.2 % of frame
area error ±1 %
01 · Segmentation

Octopus segmenter

0.642mask IoU
on held-out video

Pixel-accurate silhouette of a soft-bodied, camouflaging animal — with no prompt at inference. It scores above the 0.374 its own teacher architecture reaches zero-shot per frame, at a fraction of the size. Body area, the channel posture and masked-motion actually read, is accurate to about one percent.

LR-ASPP / MobileNetV3 3.2 M params 768² input single class CPU-capable
Colour aquarium frame showing the octopus stretched up the side of the tank. Reaching out of water Reaching0.88 Exploration0.09 No octopus0.02
02 · Behaviour

Ethogram classifier

0.665macro-F1
6 classes

Scores a 20-second clip against a six-class ethogram — with absence as one of the classes, so the number is honest. A model that always answers "no octopus" scores 0.100 here. 75.4 % accuracy against a 43.1 % majority baseline.

5-member ensemble 33.5 MB frozen backbones
Night-vision aquarium frame showing an empty tank with no animal visible. p = 0.02
Daylight aquarium frame with the octopus clearly visible against the tank. p = 0.99
03 · Presence

Presence gate

0.969AUC
temporally fused

The always-on layer that decides whether anything is worth processing. It is what makes 24/7 monitoring affordable: 60 % of video is rejected before it is ever fully decoded. 0.794 AUC on a single frame, 0.969 with temporal smoothing.

frozen CLIP + MLP probe 0.6 MB 1 fps
Colour aquarium frame of the octopus with its arms extended upward toward the top of the tank. “The octopus extends its arms upward and outward, reaching above the water surface and climbing onto the edge of the tank.”
— generated on a laptop, 3 s
04 · Language

Caption model

~3 sper clip
no GPU

A 235B teacher distilled into a 2B student you can run on a 16 GB laptop. Against the base model, similarity to the teacher rises from 0.702 to 0.834 and ROUGE-L from 0.269 to 0.455. Ships quantised and self-contained.

Qwen3-VL-2B + LoRA 4-bit · 1.7 GB offline
Aquarium frame with the octopus's skeleton extracted: mantle, head and eight arm paths each traced in a different colour with joint nodes marked. 7 arms tracked
3.68 avg / frame
05 · Kinematics

Skeleton & pose

0.539arm-tip F1
vs human labels

Silhouette to anatomical graph: mantle, head and eight arms, tracked through time with an optical-flow prior. Arm-tip speed recovers the same ordering of behaviours as the language model — without ever seeing its answer.

thinning + Dijkstra per-arm tracking precision 0.722
How it works

Spend the expensive model last.

Prompting a 235B model on every clip is not a product — one labelling run over our archive would take 372 to 1,030 hours on an API the facility does not have. So cheap gates go first, and each gate's threshold was set by a measurement, not an intuition.

01
Probe

Ten seek frames per video. Rejects 60 % with no full decode.

02
Gate

Presence probe plus absolute motion, at 1 fps.

03
Clip

20-second windows. 42.7 % still hold no animal.

04
Teacher

235B VLM, five passes, offline. $0.0006 a clip.

05
Students

Five distilled models. On site, no GPU, forever.

Bar width = share of the original 892 hours still alive at each stage. About 60 hours are ever fully decoded.

What it found

A welfare baseline
nobody had to score by hand.

3,083 clips of one animal, labelled with no human in the loop. This is the measurement welfare has always lacked — and the reason "this animal is not itself today" can become a number.

The animal keeps a clock

Visible-activity rate per clock hour · present clips ÷ all extracted windows in that hour · n = 13,342 windows

Data table
HourVisible-activity rateWindows scanned
00:00–04:590.0–4.8 %1,142
05:00–06:5910.4–12.1 %414
07:00–11:590.0–2.5 %1,546
13:00–14:5924.2–24.3 %2,043
15:00–16:5939.6–42.4 %2,575
17:00 (peak)45.3 %1,434
18:00–19:5935.9–36.5 %1,648
20:00–23:593.8–11.1 %1,992

What the animal does all day

Share of 3,081 labelled clips · six-class ethogram

Day to day this moves a lot — resting alone ranges 16–73 % across the seven analysed days, so we quote the range whenever the budget is the claim.

A measurable response to people

Two measures, two panels — never one axis carrying both scales

No human present · 532 clips Human present · 522 clips

Human presence nearly doubles movement and lifts arousal from 0.46 to 0.68 — and it holds on 7 of 7 days independently (sign test, p = 0.0078).

Open release

The labels are the asset.
So we published them.

Code under Apache-2.0, labels and benchmarks under CC-BY-4.0. We release the full vote distributions rather than majority labels, because the margin predicts human agreement and a collapsed label throws that away.

5,222
teacher-labelled clips with five-pass vote distributions
~25k
captions — five disjoint samplings per clip
969
human labels and hand-drawn masks
3
frozen, video-level benchmark suites
Where it goes next

The octopus is the hard case.

Soft-bodied, camouflaging, unpredictable, hidden half the time. A pipeline that works here transfers: the teacher relabels, the students retrain, the gates stay.

AQUACULTURE

Welfare that pays for itself

Continuous, objective monitoring for the fastest-growing food sector — where welfare and yield are the same variable.

BIODIVERSITY

Camera networks that mean something

Turn passive footage into quantified ecological signal instead of unwatched storage.

ZOOS & AQUARIA

A baseline per individual

24/7 welfare baselines, and early warning when an animal departs from its own normal.

Published work

Peer-reviewed, not just posted.

IEEE OCEANS 2026
Monterey
From Footage to Ethogram: A Deployable Pipeline for Continuous Behavioural Monitoring of a Captive Octopus
The full cascade, the frozen benchmarks, and the behavioural findings on this page. Published via IEEE Xplore.
Accepted
IEEE OCEANS 2025
DISC-O — closed-loop habitat response to an octopus's light-request gesture
Detection driving a real actuator in a real tank: AI that acts, not just annotates.
Published
IEEE ISMAR 2025
Cephalopod AR — teaching marine biology through interactive 3D ocean worlds
Our public-engagement arm, live on Google Play and Web-AR.
Published
Team

A small lab shipping real systems.

Krishna Mohan
Chief Executive Officer
Siddharth Raj
AI Researcher
Shivansh Pachnanda
AI Researcher
Prof. Wolfgang Slany
Partner & consultant
TU Graz
Get in touch

Do you have cameras and no answers?

We are looking for aquarium, hatchery and field-station partners with footage already running, and for research collaborators who want the behavioural record their video is already carrying.