tile-twinx 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,10 @@
1
+ # Python-generated files
2
+ __pycache__/
3
+ *.py[oc]
4
+ build/
5
+ dist/
6
+ wheels/
7
+ *.egg-info
8
+
9
+ # Virtual environments
10
+ .venv
@@ -0,0 +1,174 @@
1
+ Metadata-Version: 2.5
2
+ Name: tile-twinx
3
+ Version: 1.0.0
4
+ Summary: Split a raw YouTube title into artist, title and extras.
5
+ Keywords: artist,metadata,music,parser,tags,title,youtube
6
+ Classifier: Development Status :: 4 - Beta
7
+ Classifier: Environment :: Console
8
+ Classifier: Intended Audience :: Developers
9
+ Classifier: Operating System :: OS Independent
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Programming Language :: Python :: 3.12
12
+ Classifier: Programming Language :: Python :: 3.13
13
+ Classifier: Topic :: Multimedia :: Sound/Audio
14
+ Classifier: Topic :: Text Processing :: Linguistic
15
+ Classifier: Typing :: Typed
16
+ Requires-Python: >=3.12
17
+ Requires-Dist: yt-dlp>=2026.8.19
18
+ Description-Content-Type: text/markdown
19
+
20
+ # tile-twinx
21
+
22
+ Takes a raw YouTube title and splits it into `artist`, `title` and `extras`.
23
+
24
+ ```
25
+ $ uv run app.py --json "Marshmello - Alone (Official Music Video)"
26
+ {"artist": "Marshmello", "title": "Alone", "extras": ["Official Music Video"]}
27
+ ```
28
+
29
+ ## Install
30
+
31
+ ```bash
32
+ uv add tile-twinx
33
+ ```
34
+
35
+ OR
36
+
37
+ ```bash
38
+ pip install tile-twinx
39
+ ```
40
+
41
+ ```python
42
+ from Tile import parse_title
43
+
44
+ parse_title("Kesariya - Arijit Singh | T-Series | 2023")
45
+ # {'artist': 'Arijit Singh', 'title': 'Kesariya', 'extras': ['T-Series', '2023']}
46
+ ```
47
+
48
+ No dependencies, Python 3.12+.
49
+
50
+ ## The demo app
51
+
52
+ `app.py` runs a list of real YouTube titles through the parser and prints a
53
+ table. With no arguments it uses its built-in list of 29 titles:
54
+
55
+ ```bash
56
+ uv run app.py # the built-in list
57
+ uv run app.py "Alan Walker - Faded" # one title
58
+ uv run app.py --file my_titles.txt # a file, one title per line
59
+ yt-dlp --print "%(title)s" URL | uv run app.py # straight from yt-dlp
60
+ uv run app.py --random 5 # 5 random ones
61
+ uv run app.py --interactive # type titles by hand
62
+ ```
63
+
64
+ ```
65
+ ARTIST TITLE EXTRAS
66
+ -------------------------- ---------------------------------- ----------------------
67
+ Marshmello Alone Official Music Video
68
+ Kesariya Kesariya Full Song, 4K, 1080p, HD
69
+ Arijit Singh Kesariya T-Series, 2023
70
+ ...
71
+
72
+ parsed 29 titles
73
+ artist found 29/29 100%
74
+ extras found 25/29 86%
75
+ ```
76
+
77
+ Flags: `--json`, `--dict`, `--explain` (why the parser decided what it did),
78
+ `--summary`, `--colour auto|always|never`. `uv run app.py --help` lists them all.
79
+
80
+ `app.py` doubles as a smoke test: it holds 29 real titles whose expected output
81
+ is in the table above, so if a parser change breaks something it shows up
82
+ there.
83
+
84
+ ## Library
85
+
86
+ ```python
87
+ from Tile import parse_title, get_context
88
+
89
+ parse_title("Marshmello - Alone (Official Music Video)")
90
+ # {'artist': 'Marshmello', 'title': 'Alone', 'extras': ['Official Music Video']}
91
+
92
+ parse_title(raw, explain=True) # adds a "debug" list of decision notes
93
+ ctx = get_context("~/my-ignore.txt", "~/my-artists.txt") # your own data files
94
+ parse_title(raw, ctx)
95
+ ```
96
+
97
+ `Tile` also exports `build_context`, `Context`, `Segment` and `Token` for
98
+ anyone poking at the internals. `get_context` caches by mtime and reloads
99
+ automatically when a data file changes.
100
+
101
+ ## Files
102
+
103
+ | Path | Purpose |
104
+ | ------------------- | ---------------------------------------------------------------- |
105
+ | `main.py` | the whole engine plus its CLI - still the single source of truth |
106
+ | `assets/ignore.txt` | junk words/phrases to strip (`4k`, `lyric`, `official ...`) |
107
+ | `assets/artist.txt` | artist names, `Main Name \| Alias` per line |
108
+ | `Tile/__init__.py` | the public API, `from Tile import ...` |
109
+ | `app.py` | the demo app |
110
+
111
+ `main.py` stays in the repo root; the wheel ships that same file as
112
+ `Tile/main.py` so it installs inside the package, and `Tile/__init__.py`
113
+ supports both layouts. The checkout and an installed copy behave identically.
114
+
115
+ Both data files are re-read whenever they change, no restart needed. `#`
116
+ starts a comment in either one.
117
+
118
+ ## How it decides
119
+
120
+ 1. **Clean** - unescape `&`, normalise unicode, unify dashes and quotes,
121
+ split the title into chunks on `-`, `|`, `•`.
122
+ 2. **Brackets** - each `(...)` / `[...]` / `{...}` group is either metadata
123
+ (`(Official Video)`, `(feat. X)`, `(from "Album")`) or real title text
124
+ (`(Humming)`). A group whose every word is junk becomes an extra.
125
+ 3. **Junk** - `ignore.txt` words go anywhere, except the handful that also start
126
+ or end real titles (`old`, `song`, `video`, `full`, `audio`, ...), which are
127
+ only removed next to other junk. That is what keeps _Old Town Road_ and
128
+ _Video Games_ intact while _Song 4K_ still loses its junk.
129
+ 4. **Credits** - `feat./ft./prod./cover by/from "..."` are pulled out before
130
+ junk removal, then folded into the artist.
131
+ 5. **Artist** - the best `artist.txt` match: a chunk that _is_ the artist beats
132
+ one that merely contains the name, the earlier chunk beats a later one, and
133
+ within a chunk the longest match wins. It then extends over `x`, `&`, `,`
134
+ and `feat.` collaborators. Without a match it falls back to the first chunk,
135
+ then a `Song by Artist` form, then the last chunk.
136
+ 6. **Title** - whatever is left, with extras in the order they appeared.
137
+
138
+ ## Adding to the data files
139
+
140
+ ```ini
141
+ # ignore.txt - a single word is stripped anywhere, a phrase too
142
+ official teaser # safe anywhere: never part of a title
143
+
144
+ # artist.txt - aliases after a pipe, matched without case or accents
145
+ Alan Walker | Alan W.
146
+
147
+ # artist.txt - "!" marks a channel or label, never a performer
148
+ !Saregama
149
+ ```
150
+
151
+ A label in its own chunk becomes an extra, so `Kesariya - Arijit Singh |
152
+ T-Series | 2023` gives `Arijit Singh / Kesariya / ["T-Series", "2023"]`, while
153
+ `T-Series - Kesariya` still credits T-Series as the artist.
154
+
155
+ Run `uv run app.py` after editing: it reloads the data files, so a regression
156
+ shows up immediately in the table.
157
+
158
+ ## Known limits
159
+
160
+ - An unspaced dash is never a separator (`Marshmello-Alone-Official` stays one
161
+ word), because `Spider-Man` and `Merry-Go-Round` outnumber the broken form.
162
+ - When a title and its artist are both unknown _and_ both are short, the first
163
+ chunk is taken as the artist. Adding the name to `assets/artist.txt` fixes
164
+ the whole family of titles at once.
165
+
166
+ ## Publishing
167
+
168
+ ```bash
169
+ uv build # dist/*.whl + dist/*.tar.gz
170
+ twine check dist/* # verify metadata without uploading
171
+ uv publish # needs a token in ~/.pypirc or UV_PUBLISH_TOKEN
172
+ ```
173
+
174
+ Distribution name `tile-twinx`, import name `Tile`.
@@ -0,0 +1,155 @@
1
+ # tile-twinx
2
+
3
+ Takes a raw YouTube title and splits it into `artist`, `title` and `extras`.
4
+
5
+ ```
6
+ $ uv run app.py --json "Marshmello - Alone (Official Music Video)"
7
+ {"artist": "Marshmello", "title": "Alone", "extras": ["Official Music Video"]}
8
+ ```
9
+
10
+ ## Install
11
+
12
+ ```bash
13
+ uv add tile-twinx
14
+ ```
15
+
16
+ OR
17
+
18
+ ```bash
19
+ pip install tile-twinx
20
+ ```
21
+
22
+ ```python
23
+ from Tile import parse_title
24
+
25
+ parse_title("Kesariya - Arijit Singh | T-Series | 2023")
26
+ # {'artist': 'Arijit Singh', 'title': 'Kesariya', 'extras': ['T-Series', '2023']}
27
+ ```
28
+
29
+ No dependencies, Python 3.12+.
30
+
31
+ ## The demo app
32
+
33
+ `app.py` runs a list of real YouTube titles through the parser and prints a
34
+ table. With no arguments it uses its built-in list of 29 titles:
35
+
36
+ ```bash
37
+ uv run app.py # the built-in list
38
+ uv run app.py "Alan Walker - Faded" # one title
39
+ uv run app.py --file my_titles.txt # a file, one title per line
40
+ yt-dlp --print "%(title)s" URL | uv run app.py # straight from yt-dlp
41
+ uv run app.py --random 5 # 5 random ones
42
+ uv run app.py --interactive # type titles by hand
43
+ ```
44
+
45
+ ```
46
+ ARTIST TITLE EXTRAS
47
+ -------------------------- ---------------------------------- ----------------------
48
+ Marshmello Alone Official Music Video
49
+ Kesariya Kesariya Full Song, 4K, 1080p, HD
50
+ Arijit Singh Kesariya T-Series, 2023
51
+ ...
52
+
53
+ parsed 29 titles
54
+ artist found 29/29 100%
55
+ extras found 25/29 86%
56
+ ```
57
+
58
+ Flags: `--json`, `--dict`, `--explain` (why the parser decided what it did),
59
+ `--summary`, `--colour auto|always|never`. `uv run app.py --help` lists them all.
60
+
61
+ `app.py` doubles as a smoke test: it holds 29 real titles whose expected output
62
+ is in the table above, so if a parser change breaks something it shows up
63
+ there.
64
+
65
+ ## Library
66
+
67
+ ```python
68
+ from Tile import parse_title, get_context
69
+
70
+ parse_title("Marshmello - Alone (Official Music Video)")
71
+ # {'artist': 'Marshmello', 'title': 'Alone', 'extras': ['Official Music Video']}
72
+
73
+ parse_title(raw, explain=True) # adds a "debug" list of decision notes
74
+ ctx = get_context("~/my-ignore.txt", "~/my-artists.txt") # your own data files
75
+ parse_title(raw, ctx)
76
+ ```
77
+
78
+ `Tile` also exports `build_context`, `Context`, `Segment` and `Token` for
79
+ anyone poking at the internals. `get_context` caches by mtime and reloads
80
+ automatically when a data file changes.
81
+
82
+ ## Files
83
+
84
+ | Path | Purpose |
85
+ | ------------------- | ---------------------------------------------------------------- |
86
+ | `main.py` | the whole engine plus its CLI - still the single source of truth |
87
+ | `assets/ignore.txt` | junk words/phrases to strip (`4k`, `lyric`, `official ...`) |
88
+ | `assets/artist.txt` | artist names, `Main Name \| Alias` per line |
89
+ | `Tile/__init__.py` | the public API, `from Tile import ...` |
90
+ | `app.py` | the demo app |
91
+
92
+ `main.py` stays in the repo root; the wheel ships that same file as
93
+ `Tile/main.py` so it installs inside the package, and `Tile/__init__.py`
94
+ supports both layouts. The checkout and an installed copy behave identically.
95
+
96
+ Both data files are re-read whenever they change, no restart needed. `#`
97
+ starts a comment in either one.
98
+
99
+ ## How it decides
100
+
101
+ 1. **Clean** - unescape `&`, normalise unicode, unify dashes and quotes,
102
+ split the title into chunks on `-`, `|`, `•`.
103
+ 2. **Brackets** - each `(...)` / `[...]` / `{...}` group is either metadata
104
+ (`(Official Video)`, `(feat. X)`, `(from "Album")`) or real title text
105
+ (`(Humming)`). A group whose every word is junk becomes an extra.
106
+ 3. **Junk** - `ignore.txt` words go anywhere, except the handful that also start
107
+ or end real titles (`old`, `song`, `video`, `full`, `audio`, ...), which are
108
+ only removed next to other junk. That is what keeps _Old Town Road_ and
109
+ _Video Games_ intact while _Song 4K_ still loses its junk.
110
+ 4. **Credits** - `feat./ft./prod./cover by/from "..."` are pulled out before
111
+ junk removal, then folded into the artist.
112
+ 5. **Artist** - the best `artist.txt` match: a chunk that _is_ the artist beats
113
+ one that merely contains the name, the earlier chunk beats a later one, and
114
+ within a chunk the longest match wins. It then extends over `x`, `&`, `,`
115
+ and `feat.` collaborators. Without a match it falls back to the first chunk,
116
+ then a `Song by Artist` form, then the last chunk.
117
+ 6. **Title** - whatever is left, with extras in the order they appeared.
118
+
119
+ ## Adding to the data files
120
+
121
+ ```ini
122
+ # ignore.txt - a single word is stripped anywhere, a phrase too
123
+ official teaser # safe anywhere: never part of a title
124
+
125
+ # artist.txt - aliases after a pipe, matched without case or accents
126
+ Alan Walker | Alan W.
127
+
128
+ # artist.txt - "!" marks a channel or label, never a performer
129
+ !Saregama
130
+ ```
131
+
132
+ A label in its own chunk becomes an extra, so `Kesariya - Arijit Singh |
133
+ T-Series | 2023` gives `Arijit Singh / Kesariya / ["T-Series", "2023"]`, while
134
+ `T-Series - Kesariya` still credits T-Series as the artist.
135
+
136
+ Run `uv run app.py` after editing: it reloads the data files, so a regression
137
+ shows up immediately in the table.
138
+
139
+ ## Known limits
140
+
141
+ - An unspaced dash is never a separator (`Marshmello-Alone-Official` stays one
142
+ word), because `Spider-Man` and `Merry-Go-Round` outnumber the broken form.
143
+ - When a title and its artist are both unknown _and_ both are short, the first
144
+ chunk is taken as the artist. Adding the name to `assets/artist.txt` fixes
145
+ the whole family of titles at once.
146
+
147
+ ## Publishing
148
+
149
+ ```bash
150
+ uv build # dist/*.whl + dist/*.tar.gz
151
+ twine check dist/* # verify metadata without uploading
152
+ uv publish # needs a token in ~/.pypirc or UV_PUBLISH_TOKEN
153
+ ```
154
+
155
+ Distribution name `tile-twinx`, import name `Tile`.
@@ -0,0 +1,67 @@
1
+ """Tile - split a raw YouTube title into artist / title / extras.
2
+
3
+ from Tile import parse_title
4
+
5
+ parse_title("Marshmello - Alone (Official Music Video)")
6
+ # {'artist': 'Marshmello', 'title': 'Alone', 'extras': ['Official Music Video']}
7
+
8
+ The engine lives in ``main.py`` at the repository root, and the wheel ships
9
+ that same file as ``Tile/main.py`` so it installs inside the package. Both
10
+ layouts are supported, so the checkout and an installed copy behave the same:
11
+
12
+ checkout Tile/__init__.py + ../main.py + ../assets/
13
+ installed Tile/__init__.py + Tile/main.py + Tile/assets/
14
+
15
+ Swap in your own data files with ``get_context``:
16
+
17
+ from Tile import get_context, parse_title
18
+
19
+ ctx = get_context("~/my-ignore.txt", "~/my-artists.txt")
20
+ parse_title(raw, ctx)
21
+
22
+ A context is cached and reloaded whenever either file's mtime changes, so
23
+ editing ``assets/artist.txt`` takes effect immediately.
24
+ """
25
+
26
+ from __future__ import annotations
27
+
28
+ import sys
29
+ from pathlib import Path
30
+
31
+ _HERE = Path(__file__).resolve().parent
32
+
33
+ if (_HERE / "main.py").exists():
34
+ from . import main as _engine # installed wheel
35
+ else:
36
+ sys.path.insert(0, str(_HERE.parent)) # repo checkout
37
+ import main as _engine # type: ignore[no-redef]
38
+ # pretend the root file is the submodule, so `from Tile.main import ...`
39
+ # works in the checkout exactly as it does once installed
40
+ sys.modules[f"{__name__}.main"] = _engine
41
+
42
+ __version__ = "1.0.0"
43
+
44
+ # The engine module itself, reachable as `Tile.main` in both layouts: the real
45
+ # submodule once installed, the repo's root main.py in a checkout.
46
+ main = _engine
47
+
48
+ DEFAULT_IGNORE_FILE = _engine.DEFAULT_IGNORE_FILE
49
+ DEFAULT_ARTIST_FILE = _engine.DEFAULT_ARTIST_FILE
50
+ Context = _engine.Context
51
+ Segment = _engine.Segment
52
+ Token = _engine.Token
53
+ build_context = _engine.build_context
54
+ get_context = _engine.get_context
55
+ parse_title = _engine.parse_title
56
+
57
+ __all__ = [
58
+ "DEFAULT_ARTIST_FILE",
59
+ "DEFAULT_IGNORE_FILE",
60
+ "Context",
61
+ "Segment",
62
+ "Token",
63
+ "build_context",
64
+ "get_context",
65
+ "parse_title",
66
+ "__version__",
67
+ ]