video-performance-analyzer 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- video_performance_analyzer-0.1.0/.github/workflows/publish.yml +62 -0
- video_performance_analyzer-0.1.0/.github/workflows/test.yml +31 -0
- video_performance_analyzer-0.1.0/.gitignore +13 -0
- video_performance_analyzer-0.1.0/CONTRIBUTING.md +108 -0
- video_performance_analyzer-0.1.0/LICENSE +36 -0
- video_performance_analyzer-0.1.0/PKG-INFO +253 -0
- video_performance_analyzer-0.1.0/README.md +222 -0
- video_performance_analyzer-0.1.0/RELEASING.md +60 -0
- video_performance_analyzer-0.1.0/pyproject.toml +76 -0
- video_performance_analyzer-0.1.0/tests/test_analysis.py +111 -0
- video_performance_analyzer-0.1.0/tests/test_explain.py +29 -0
- video_performance_analyzer-0.1.0/tests/test_quiet.py +96 -0
- video_performance_analyzer-0.1.0/tests/test_storage.py +124 -0
- video_performance_analyzer-0.1.0/vpa/__init__.py +3 -0
- video_performance_analyzer-0.1.0/vpa/analysis/__init__.py +1 -0
- video_performance_analyzer-0.1.0/vpa/analysis/compare.py +192 -0
- video_performance_analyzer-0.1.0/vpa/analysis/features.py +170 -0
- video_performance_analyzer-0.1.0/vpa/cli.py +699 -0
- video_performance_analyzer-0.1.0/vpa/config.py +164 -0
- video_performance_analyzer-0.1.0/vpa/db.py +382 -0
- video_performance_analyzer-0.1.0/vpa/explain.py +152 -0
- video_performance_analyzer-0.1.0/vpa/learn.py +231 -0
- video_performance_analyzer-0.1.0/vpa/metrics.py +171 -0
- video_performance_analyzer-0.1.0/vpa/pipeline.py +242 -0
- video_performance_analyzer-0.1.0/vpa/providers/__init__.py +26 -0
- video_performance_analyzer-0.1.0/vpa/providers/base.py +29 -0
- video_performance_analyzer-0.1.0/vpa/providers/openai_compatible.py +111 -0
- video_performance_analyzer-0.1.0/vpa/quiet.py +175 -0
- video_performance_analyzer-0.1.0/vpa/recommend.py +241 -0
- video_performance_analyzer-0.1.0/vpa/report.py +278 -0
- video_performance_analyzer-0.1.0/vpa/tribe.py +370 -0
- video_performance_analyzer-0.1.0/vpa/tui.py +380 -0
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
name: publish
|
|
2
|
+
|
|
3
|
+
# Publishes to PyPI when you create a GitHub Release.
|
|
4
|
+
#
|
|
5
|
+
# Uses Trusted Publishing (OIDC): PyPI verifies the release came from this
|
|
6
|
+
# repository's workflow, so there is no API token to store, leak or rotate.
|
|
7
|
+
# One-time setup is described in RELEASING.md.
|
|
8
|
+
|
|
9
|
+
on:
|
|
10
|
+
release:
|
|
11
|
+
types: [published]
|
|
12
|
+
workflow_dispatch: # lets you do a TestPyPI dry run by hand
|
|
13
|
+
|
|
14
|
+
jobs:
|
|
15
|
+
build:
|
|
16
|
+
runs-on: ubuntu-latest
|
|
17
|
+
steps:
|
|
18
|
+
- uses: actions/checkout@v4
|
|
19
|
+
- uses: actions/setup-python@v5
|
|
20
|
+
with:
|
|
21
|
+
python-version: "3.12"
|
|
22
|
+
|
|
23
|
+
- name: Verify the version tag matches the package version
|
|
24
|
+
# A release tagged v0.2.0 that ships 0.1.0 is the classic way to publish
|
|
25
|
+
# something nobody can install correctly. Fail loudly instead.
|
|
26
|
+
if: github.event_name == 'release'
|
|
27
|
+
run: |
|
|
28
|
+
TAG="${GITHUB_REF_NAME#v}"
|
|
29
|
+
PKG=$(grep -m1 '^version' pyproject.toml | cut -d'"' -f2)
|
|
30
|
+
echo "tag=$TAG package=$PKG"
|
|
31
|
+
[ "$TAG" = "$PKG" ] || { echo "::error::tag $TAG != version $PKG"; exit 1; }
|
|
32
|
+
|
|
33
|
+
- name: Run the tests before shipping anything
|
|
34
|
+
run: |
|
|
35
|
+
python -m pip install --upgrade pip
|
|
36
|
+
pip install -e '.[dev]'
|
|
37
|
+
ruff check vpa tests
|
|
38
|
+
pytest -q
|
|
39
|
+
|
|
40
|
+
- name: Build
|
|
41
|
+
run: |
|
|
42
|
+
pip install build twine
|
|
43
|
+
python -m build
|
|
44
|
+
twine check dist/*
|
|
45
|
+
|
|
46
|
+
- uses: actions/upload-artifact@v4
|
|
47
|
+
with:
|
|
48
|
+
name: dist
|
|
49
|
+
path: dist/
|
|
50
|
+
|
|
51
|
+
publish:
|
|
52
|
+
needs: build
|
|
53
|
+
runs-on: ubuntu-latest
|
|
54
|
+
environment: pypi # add required reviewers here if you want a gate
|
|
55
|
+
permissions:
|
|
56
|
+
id-token: write # required for Trusted Publishing
|
|
57
|
+
steps:
|
|
58
|
+
- uses: actions/download-artifact@v4
|
|
59
|
+
with:
|
|
60
|
+
name: dist
|
|
61
|
+
path: dist/
|
|
62
|
+
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
name: test
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches: [main]
|
|
6
|
+
pull_request:
|
|
7
|
+
|
|
8
|
+
jobs:
|
|
9
|
+
test:
|
|
10
|
+
runs-on: ${{ matrix.os }}
|
|
11
|
+
strategy:
|
|
12
|
+
fail-fast: false
|
|
13
|
+
matrix:
|
|
14
|
+
os: [ubuntu-latest, macos-latest]
|
|
15
|
+
python: ["3.11", "3.12", "3.13"]
|
|
16
|
+
steps:
|
|
17
|
+
- uses: actions/checkout@v4
|
|
18
|
+
- uses: actions/setup-python@v5
|
|
19
|
+
with:
|
|
20
|
+
python-version: ${{ matrix.python }}
|
|
21
|
+
- name: Install
|
|
22
|
+
# A clean install on purpose: the scoring extras are NOT installed here,
|
|
23
|
+
# so anything the core accidentally depends on will fail this job. That
|
|
24
|
+
# is exactly how an undeclared tqdm dependency slipped through once.
|
|
25
|
+
run: |
|
|
26
|
+
python -m pip install --upgrade pip
|
|
27
|
+
pip install -e '.[dev]'
|
|
28
|
+
- name: Lint
|
|
29
|
+
run: ruff check vpa tests
|
|
30
|
+
- name: Test
|
|
31
|
+
run: pytest -q
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Thanks for looking. This is a small project with a specific point of view, so
|
|
4
|
+
this file covers what that is — it should save you guessing.
|
|
5
|
+
|
|
6
|
+
## Setup
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
git clone https://github.com/christopher-inegbedion/video-performance-analyzer.git
|
|
10
|
+
cd video-performance-analyzer
|
|
11
|
+
python3.12 -m venv .venv && source .venv/bin/activate
|
|
12
|
+
pip install -e '.[dev]'
|
|
13
|
+
|
|
14
|
+
pytest -q
|
|
15
|
+
ruff check vpa tests
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
That is enough for everything except actually scoring a video. For that you also
|
|
19
|
+
need ffmpeg, roughly 3GB of model weights and a lot of patience:
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install -e '.[tribe]'
|
|
23
|
+
pip install git+https://github.com/facebookresearch/tribev2.git
|
|
24
|
+
vpa doctor # says exactly what is missing
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
**CI deliberately does not install the scoring extras.** A clean core install is
|
|
28
|
+
what catches undeclared dependencies — that is how an accidental reliance on
|
|
29
|
+
`tqdm` was found, after it passed locally for days. If you add something to
|
|
30
|
+
`vpa/` that imports a new package, declare it in `pyproject.toml`.
|
|
31
|
+
|
|
32
|
+
## The one rule that matters
|
|
33
|
+
|
|
34
|
+
**Never show a number without being able to say what it means and what it does
|
|
35
|
+
not.**
|
|
36
|
+
|
|
37
|
+
TRIBE v2 predicts fMRI response. It does not predict attention, watch time,
|
|
38
|
+
clicks or sales. The whole tool is built around not letting that slide:
|
|
39
|
+
|
|
40
|
+
- every metric has an entry in `vpa/explain.py`, and a test asserts it
|
|
41
|
+
- the same caveats are injected into the LLM's context, so recommendations
|
|
42
|
+
inherit them rather than overclaiming
|
|
43
|
+
- known artefacts are handled, not reported as insight — a cut to black spikes
|
|
44
|
+
the curve every time, so the tail is excluded from peak detection
|
|
45
|
+
|
|
46
|
+
If you add a metric, add its explanation in the same change. `tests/test_explain.py`
|
|
47
|
+
will fail if you do not.
|
|
48
|
+
|
|
49
|
+
## Where things live
|
|
50
|
+
|
|
51
|
+
| Module | Responsibility |
|
|
52
|
+
|---|---|
|
|
53
|
+
| `vpa/tribe.py` | Runs TRIBE and patches what stops it working off a GPU cluster |
|
|
54
|
+
| `vpa/analysis/` | Predictions → response curve → comparisons |
|
|
55
|
+
| `vpa/learn.py` | Correlating predicted features with real published performance |
|
|
56
|
+
| `vpa/recommend.py` | Prompt construction; the rules the model must obey |
|
|
57
|
+
| `vpa/providers/` | LLM backends |
|
|
58
|
+
| `vpa/explain.py` | What everything means, and its limits |
|
|
59
|
+
| `vpa/quiet.py` | Keeping third-party output from wrecking the progress display |
|
|
60
|
+
|
|
61
|
+
## Good first contributions
|
|
62
|
+
|
|
63
|
+
**Platform connectors for metrics.** Performance is entered by hand or via CSV
|
|
64
|
+
today. `vpa/metrics.py` has a documented seam: anything that can produce
|
|
65
|
+
`{video, views, likes, ...}` rows can feed `record()`. Instagram Graph API,
|
|
66
|
+
TikTok, YouTube Analytics are all wanted. Keep manual entry working — it is the
|
|
67
|
+
only thing that covers every platform.
|
|
68
|
+
|
|
69
|
+
**LLM backends.** `vpa/providers/` has a `Provider` protocol and one
|
|
70
|
+
OpenAI-compatible implementation covering most services. Anything with a
|
|
71
|
+
different shape (Anthropic's native API, Bedrock, Vertex) needs its own class
|
|
72
|
+
registered in `get_provider`.
|
|
73
|
+
|
|
74
|
+
**Speed.** Scoring runs at roughly 96x realtime on CPU. Lowering `target_fps`
|
|
75
|
+
helps proportionally, but nobody has measured what it costs in accuracy. A
|
|
76
|
+
careful comparison of the same video at 24 and 12 fps would be genuinely
|
|
77
|
+
valuable and needs no new code.
|
|
78
|
+
|
|
79
|
+
**Honesty about small samples.** `vpa/learn.py` refuses to call correlations
|
|
80
|
+
findings below five labelled videos. That threshold is a guess. If you know the
|
|
81
|
+
statistics better than I do, improve it.
|
|
82
|
+
|
|
83
|
+
## Style
|
|
84
|
+
|
|
85
|
+
- Comments explain *why*, especially where the code looks odd. Most of
|
|
86
|
+
`vpa/tribe.py` is strange for a reason and each reason is written down.
|
|
87
|
+
- Errors should say what to do next. `vpa doctor` is the model: state what is
|
|
88
|
+
missing and give the command that fixes it.
|
|
89
|
+
- Prefer failing loudly over guessing quietly. A silent wrong number is worse
|
|
90
|
+
than an error.
|
|
91
|
+
- `ruff check vpa tests` must pass. Line length 100.
|
|
92
|
+
|
|
93
|
+
## Testing
|
|
94
|
+
|
|
95
|
+
`pytest -q` needs no models, no API key and no network. Anything requiring those
|
|
96
|
+
belongs behind a skip.
|
|
97
|
+
|
|
98
|
+
Test behaviour rather than implementation. The most valuable test in the suite
|
|
99
|
+
checks that the progress relay reports *climbing* values — it caught a bug where
|
|
100
|
+
every relayed value was zero, which would have shipped a frozen progress bar as
|
|
101
|
+
a fix for a frozen progress bar.
|
|
102
|
+
|
|
103
|
+
## Licences
|
|
104
|
+
|
|
105
|
+
This tool is MIT. The models are not. TRIBE v2 is CC BY-NC 4.0 — non-commercial
|
|
106
|
+
— and the language pathway needs a gated Llama repo. Please do not add anything
|
|
107
|
+
that obscures those terms from users; the current behaviour of stating them
|
|
108
|
+
plainly is deliberate.
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 video-performance-analyzer contributors
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
NOTE ON MODEL LICENCES
|
|
26
|
+
|
|
27
|
+
This tool orchestrates third-party models that carry their own terms:
|
|
28
|
+
|
|
29
|
+
* Meta TRIBE v2 (facebook/tribev2, facebook/tribev2-mini) is released under
|
|
30
|
+
CC BY-NC 4.0 — NON-COMMERCIAL use only. Using it to optimise commercial
|
|
31
|
+
advertising is arguably outside that licence. This is your decision to make,
|
|
32
|
+
not this tool's.
|
|
33
|
+
* TRIBE's language pathway uses meta-llama/Llama-3.2-1B, a gated repository
|
|
34
|
+
requiring licence acceptance and a HuggingFace token.
|
|
35
|
+
|
|
36
|
+
The MIT licence above covers this tool's own source code only.
|
|
@@ -0,0 +1,253 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: video-performance-analyzer
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Predict and compare how videos land, using Meta's TRIBE v2 brain-encoding model, with LLM recommendations that learn from real performance.
|
|
5
|
+
Project-URL: Homepage, https://github.com/christopher-inegbedion/video-performance-analyzer
|
|
6
|
+
Project-URL: Issues, https://github.com/christopher-inegbedion/video-performance-analyzer/issues
|
|
7
|
+
Author: video-performance-analyzer contributors
|
|
8
|
+
License: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: analysis,creative,ffmpeg,fmri,neuroscience,tribe,video
|
|
11
|
+
Classifier: Development Status :: 3 - Alpha
|
|
12
|
+
Classifier: Environment :: Console
|
|
13
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
14
|
+
Classifier: Programming Language :: Python :: 3
|
|
15
|
+
Classifier: Topic :: Multimedia :: Video
|
|
16
|
+
Requires-Python: >=3.11
|
|
17
|
+
Requires-Dist: httpx>=0.27
|
|
18
|
+
Requires-Dist: numpy>=1.26
|
|
19
|
+
Requires-Dist: platformdirs>=4.2
|
|
20
|
+
Requires-Dist: rich>=13.7
|
|
21
|
+
Requires-Dist: tomli-w>=1.0
|
|
22
|
+
Requires-Dist: tqdm>=4.66
|
|
23
|
+
Requires-Dist: typer>=0.12
|
|
24
|
+
Provides-Extra: dev
|
|
25
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
26
|
+
Requires-Dist: ruff>=0.5; extra == 'dev'
|
|
27
|
+
Provides-Extra: tribe
|
|
28
|
+
Requires-Dist: faster-whisper>=1.0; extra == 'tribe'
|
|
29
|
+
Requires-Dist: torch<2.7,>=2.5; extra == 'tribe'
|
|
30
|
+
Description-Content-Type: text/markdown
|
|
31
|
+
|
|
32
|
+
# video-performance-analyzer
|
|
33
|
+
|
|
34
|
+
[](https://pypi.org/project/video-performance-analyzer/)
|
|
35
|
+
[](https://pypi.org/project/video-performance-analyzer/)
|
|
36
|
+
[](https://github.com/christopher-inegbedion/video-performance-analyzer/actions/workflows/test.yml)
|
|
37
|
+
[](LICENSE)
|
|
38
|
+
|
|
39
|
+
Predict how a video lands before you publish it — then find out whether the
|
|
40
|
+
prediction was right, and get better advice because of it.
|
|
41
|
+
|
|
42
|
+
`vpa` runs Meta's **TRIBE v2** brain-encoding model over your video, reduces the
|
|
43
|
+
output to a response curve you can read, compares it against a reference video
|
|
44
|
+
and against your own back catalogue, and asks an LLM for concrete editing
|
|
45
|
+
recommendations. When you later record what the video actually did — views,
|
|
46
|
+
likes, watch-through — it learns, and goes back to revise advice it has already
|
|
47
|
+
given you.
|
|
48
|
+
|
|
49
|
+
```
|
|
50
|
+
vpa analyse my-cut-v3.mp4 --reference competitor-ad.mp4 \
|
|
51
|
+
--segments "hook:0-3,montage:3-16,card:16-19,end:19-24"
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
```
|
|
55
|
+
╭─ my-cut-v3 ──────────────────────────────────────────────────╮
|
|
56
|
+
│ 24.1s · 25 timesteps · Audio, Video │
|
|
57
|
+
╰──────────────────────────────────────────────────────────────╯
|
|
58
|
+
╭─ predicted response over time ───────────────────────────────╮
|
|
59
|
+
│ ▃▄▄▅▅▆▇▇▆▇▇▇▆▅▄▃▃▃▃▂▂▂▂█ │
|
|
60
|
+
│ 0s 24s │
|
|
61
|
+
╰──────────────────────────────────────────────────────────────╯
|
|
62
|
+
opening_2s 0.882 Mean response over the first two seconds.
|
|
63
|
+
trough_value 0.777 Lowest normalised response.
|
|
64
|
+
decay +0.055 Opening minus closing.
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
69
|
+
## What this actually measures — read this first
|
|
70
|
+
|
|
71
|
+
TRIBE v2 predicts **fMRI brain response**. It does not predict attention, watch
|
|
72
|
+
time, clicks, or sales. Treating "more predicted response" as "better advert" is
|
|
73
|
+
an inference the model does not make and its authors do not claim.
|
|
74
|
+
|
|
75
|
+
This matters enough that the tool is built around it:
|
|
76
|
+
|
|
77
|
+
- Every metric shown has a plain-English explanation attached (`vpa explain`).
|
|
78
|
+
- The same caveats are injected into the LLM's context, so recommendations
|
|
79
|
+
inherit them rather than overclaiming.
|
|
80
|
+
- Known artefacts are flagged automatically. A cut to black spikes the curve
|
|
81
|
+
every time; `vpa` detects that and excludes it from peak-finding rather than
|
|
82
|
+
reporting it as your ending landing.
|
|
83
|
+
- Position matters more than content. In testing, the *same images* scored
|
|
84
|
+
highest in one edit and lowest in another purely because of where they sat.
|
|
85
|
+
|
|
86
|
+
Use it to rank variants of the same idea. Don't use it as a verdict on one video.
|
|
87
|
+
|
|
88
|
+
## Install
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
pip install video-performance-analyzer # core
|
|
92
|
+
pip install 'video-performance-analyzer[tribe]' # + scoring dependencies (~3GB of models)
|
|
93
|
+
pip install git+https://github.com/facebookresearch/tribev2.git
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
TRIBE itself is not on PyPI, so that last line is always needed for scoring.
|
|
97
|
+
|
|
98
|
+
<details>
|
|
99
|
+
<summary>From source instead</summary>
|
|
100
|
+
|
|
101
|
+
```bash
|
|
102
|
+
git clone https://github.com/christopher-inegbedion/video-performance-analyzer.git
|
|
103
|
+
cd video-performance-analyzer
|
|
104
|
+
python3.12 -m venv .venv && source .venv/bin/activate
|
|
105
|
+
pip install -e '.[dev]'
|
|
106
|
+
```
|
|
107
|
+
</details>
|
|
108
|
+
|
|
109
|
+
You also need **ffmpeg** (`brew install ffmpeg` / `apt install ffmpeg`) and an
|
|
110
|
+
API key for whichever LLM you point it at:
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
export OPENROUTER_API_KEY=sk-or-...
|
|
114
|
+
vpa doctor # tells you exactly what is missing
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
## Use
|
|
118
|
+
|
|
119
|
+
```bash
|
|
120
|
+
vpa analyse cut.mp4 # score one video
|
|
121
|
+
vpa analyse cut.mp4 -r reference.mp4 # compare against something
|
|
122
|
+
vpa show eval_a1b2c3 # revisit a past evaluation
|
|
123
|
+
vpa list evals # everything you have run
|
|
124
|
+
vpa ask eval_a1b2c3 # ask questions about it
|
|
125
|
+
vpa explain artefacts # what the numbers can't tell you
|
|
126
|
+
vpa tui # interactive session
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
### Describing your structure
|
|
130
|
+
|
|
131
|
+
Even segments are a poor guide. Tell it where your real beats are and the report
|
|
132
|
+
speaks your language:
|
|
133
|
+
|
|
134
|
+
```bash
|
|
135
|
+
vpa analyse cut.mp4 --segments "hook:0-3,montage:3-16,card:16-19,logo:19-24"
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
### The learning loop
|
|
139
|
+
|
|
140
|
+
This is what makes the tool improve. Record what a published video did:
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
vpa metrics add my-cut-v3 --views 12400 --likes 380 --platform instagram
|
|
144
|
+
vpa metrics template -o metrics.csv && vpa metrics import metrics.csv
|
|
145
|
+
vpa metrics show # what it has learned so far
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
Once outcomes exist, two things change:
|
|
149
|
+
|
|
150
|
+
1. New recommendations weight **real performance above predicted response**.
|
|
151
|
+
2. Old evaluations become *stale* — their advice predates what you now know.
|
|
152
|
+
`vpa retrofit` revisits them and writes a new generation of recommendations
|
|
153
|
+
saying what changed. Nothing is overwritten; you can read every generation
|
|
154
|
+
with `vpa show <id> -g 1`.
|
|
155
|
+
|
|
156
|
+
The tool is deliberately honest about sample size. Below five labelled videos it
|
|
157
|
+
refuses to call correlations findings and says so in plain terms.
|
|
158
|
+
|
|
159
|
+
## Configuration
|
|
160
|
+
|
|
161
|
+
```bash
|
|
162
|
+
vpa config init # writes a commented config file
|
|
163
|
+
vpa config show # effective settings and where they came from
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
Any OpenAI-compatible endpoint works — OpenRouter (default), OpenAI, Together,
|
|
167
|
+
Groq, or a local model:
|
|
168
|
+
|
|
169
|
+
```toml
|
|
170
|
+
[llm]
|
|
171
|
+
model = "google/gemini-2.5-flash"
|
|
172
|
+
base_url = "https://openrouter.ai/api/v1"
|
|
173
|
+
|
|
174
|
+
[tribe]
|
|
175
|
+
target_fps = 24 # 60fps costs ~2.5x for identical footage
|
|
176
|
+
enable_language = false # true needs a gated Llama repo + HF token
|
|
177
|
+
|
|
178
|
+
[analysis]
|
|
179
|
+
ignore_tail_s = 1.5 # don't let the cut-to-black spike become a "finding"
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
For a fully local setup, point `base_url` at Ollama (`http://localhost:11434/v1`)
|
|
183
|
+
and no API key is needed.
|
|
184
|
+
|
|
185
|
+
## Running TRIBE on a laptop
|
|
186
|
+
|
|
187
|
+
TRIBE ships configured for Meta's GPU cluster. `vpa` patches the known problems
|
|
188
|
+
automatically and tells you what it changed:
|
|
189
|
+
|
|
190
|
+
| Problem | What happens without the fix |
|
|
191
|
+
|---|---|
|
|
192
|
+
| `device: cuda` baked into the checkpoint | `Torch not compiled with CUDA enabled` |
|
|
193
|
+
| `num_workers: 20` baked in | DataLoader workers die silently; the process sits at 3% CPU looking alive |
|
|
194
|
+
| `compute_type` hardcoded to `float16` | CPU speech extraction always fails |
|
|
195
|
+
| `uvx whisperx` resolves a broken torchaudio | `module 'torchaudio' has no attribute 'list_audio_backends'` |
|
|
196
|
+
| Language pathway needs gated `meta-llama/Llama-3.2-1B` | 401 on an otherwise working run |
|
|
197
|
+
|
|
198
|
+
The last two are why word timings are generated locally with faster-whisper and
|
|
199
|
+
written to the `.tsv` cache TRIBE reads — whisperx is never invoked. The language
|
|
200
|
+
pathway is **off by default**, so vision + audition work with no licence gate.
|
|
201
|
+
|
|
202
|
+
### How long it takes
|
|
203
|
+
|
|
204
|
+
Scoring is dominated entirely by the video encoder. Setup — importing the
|
|
205
|
+
package, loading the checkpoint, building events — is about 8 seconds. Everything
|
|
206
|
+
after that scales with how much footage you feed it.
|
|
207
|
+
|
|
208
|
+
Measured on an Apple Silicon laptop (CPU only), roughly **96x realtime**:
|
|
209
|
+
|
|
210
|
+
| video length | time to score |
|
|
211
|
+
|---:|---:|
|
|
212
|
+
| 10s | ~16 min |
|
|
213
|
+
| 15s | ~24 min |
|
|
214
|
+
| 24s | ~38 min |
|
|
215
|
+
| 60s | ~96 min |
|
|
216
|
+
|
|
217
|
+
**This is not an interactive tool.** Start a run and come back to it.
|
|
218
|
+
|
|
219
|
+
Three ways to make it tractable:
|
|
220
|
+
|
|
221
|
+
- **Halve the frame rate.** `target_fps = 12` roughly halves the encode, because
|
|
222
|
+
cost is proportional to frames. V-JEPA samples frames rather than reading every
|
|
223
|
+
one, so the quality cost is plausibly small — but that is untested, so measure
|
|
224
|
+
it on your own material before trusting it.
|
|
225
|
+
- **Score an excerpt.** For a feed asset the first 6-10 seconds is where the
|
|
226
|
+
scroll decision happens. Scoring only the opening is a legitimate strategy.
|
|
227
|
+
- **Use a GPU.** This is what the model was built for and it is a different order
|
|
228
|
+
of magnitude. Set `device = "cuda"`.
|
|
229
|
+
|
|
230
|
+
There is no caching or warm-start trick that helps: the cost is the encoder, not
|
|
231
|
+
the setup.
|
|
232
|
+
|
|
233
|
+
## Licences
|
|
234
|
+
|
|
235
|
+
This tool is MIT. The models are not:
|
|
236
|
+
|
|
237
|
+
- **TRIBE v2** is **CC BY-NC 4.0 — non-commercial**. Using it to optimise
|
|
238
|
+
commercial advertising is arguably outside that licence. That is your call to
|
|
239
|
+
make deliberately, and the tool says so rather than hiding it.
|
|
240
|
+
- The language pathway uses **Llama-3.2-1B**, which is gated and carries Meta's
|
|
241
|
+
own terms.
|
|
242
|
+
|
|
243
|
+
## Development
|
|
244
|
+
|
|
245
|
+
```bash
|
|
246
|
+
pip install -e '.[dev]'
|
|
247
|
+
pytest -q
|
|
248
|
+
ruff check vpa
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
Contributions welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). Platform
|
|
252
|
+
connectors for metrics ingestion (`vpa/metrics.py` has a documented seam) and
|
|
253
|
+
additional LLM providers (`vpa/providers/`) are the most useful places to start.
|