sedd 0.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
sedd-0.0.0/LICENSE ADDED
@@ -0,0 +1,32 @@
1
+ Note that the MIT license only applies to the software. Any downloaded
2
+ or generated data dumps are under various versions of CC-By-SA:
3
+
4
+ * CC-By-SA 2.5: https://creativecommons.org/licenses/by-sa/2.5/
5
+ * CC-By-SA 3.0: https://creativecommons.org/licenses/by-sa/3.0/
6
+ * CC-By-SA 4.0: https://creativecommons.org/licenses/by-sa/4.0/
7
+
8
+ For more details, and date ranges for the licenses, see
9
+ https://stackoverflow.com/help/licensing
10
+
11
+ ---
12
+
13
+ Copyright © 2024 Olivia
14
+
15
+ Permission is hereby granted, free of charge, to any person obtaining
16
+ a copy of this software and associated documentation files (the "Software"),
17
+ to deal in the Software without restriction, including without limitation
18
+ the rights to use, copy, modify, merge, publish, distribute, sublicense,
19
+ and/or sell copies of the Software, and to permit persons to whom the
20
+ Software is furnished to do so, subject to the following conditions:
21
+
22
+ The above copyright notice and this permission notice shall be included
23
+ in all copies or substantial portions of the Software.
24
+
25
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
26
+ EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES
27
+ OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
28
+ IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM,
29
+ DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT,
30
+ TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE
31
+ OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
32
+
sedd-0.0.0/PKG-INFO ADDED
@@ -0,0 +1,91 @@
1
+ Metadata-Version: 2.4
2
+ Name: sedd
3
+ Version: 0.0.0
4
+ Summary: Unofficial, community-made tool for downloading the Stack Exchange data dumps
5
+ Author-email: LunarWatcher <oliviawolfie@pm.me>
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/LunarWatcher/se-data-dump-transformer
8
+ Project-URL: Documentation, https://github.com/LunarWatcher/se-data-dump-transformer
9
+ Project-URL: Repository, https://github.com/LunarWatcher/se-data-dump-transformer.git
10
+ Project-URL: Issues, https://github.com/LunarWatcher/se-data-dump-transformer/issues
11
+ Project-URL: Changelog, https://github.com/LunarWatcher/se-data-dump-transformer/releases
12
+ Keywords: Stack Exchange,data dump,Stack Exchange data dump,downloader
13
+ Classifier: Development Status :: 5 - Production/Stable
14
+ Classifier: Environment :: Console
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Topic :: System :: Archiving
17
+ Classifier: Topic :: Utilities
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3.10
20
+ Classifier: Programming Language :: Python :: 3.11
21
+ Classifier: Programming Language :: Python :: 3.12
22
+ Classifier: Operating System :: POSIX :: Linux
23
+ Classifier: Operating System :: MacOS
24
+ Classifier: Operating System :: Microsoft :: Windows
25
+ Requires-Python: >=3.10
26
+ Description-Content-Type: text/markdown
27
+ License-File: LICENSE
28
+ Requires-Dist: selenium==4.32.0
29
+ Requires-Dist: undetected-geckodriver-lw>=2.1.0
30
+ Requires-Dist: desktop-notifier==5.0.1
31
+ Requires-Dist: watchdog==4.0.2
32
+ Requires-Dist: requests
33
+ Dynamic: license-file
34
+
35
+ # SE Data Dump Downloader
36
+
37
+
38
+ For more comprehensive information, please read the [main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master) on GitHub. This README contains an abridged version of the main README specifically aimed at Pypi users.
39
+
40
+ For usage problems not listed in this readme, see the main README. If no information exists, please open an issue on GitHub - keeping the tool accessible to everyone is a priority.
41
+
42
+ ---
43
+
44
+ The SE Data Dump Downloader (abbreviated `sedd`) is a command line Selenium-based utility for downloading the entire Stack Exchange data dump in their new [anti-community format](https://stackoverflow.com/help/data-dumps), since they decided not to bother providing an official "download all" button. It's one of two components that operate on the data dump in the second project, the other being the (non-python-based) SE data dump transformer - a project that converts the data dump from the not-so-useful official `.xml` format to some other formats. The pypi package is exclusively for the downloader, and does not ship with a copy of the transformer. See the main README if you're looking for the transformer.
45
+
46
+
47
+ For the pypi version, you can download it with:
48
+ ```python3
49
+ pip3 install sedd
50
+ ```
51
+
52
+ Note that there are some additional steps before you can start using it, that are detailed in this README.
53
+
54
+ ## Configuration
55
+
56
+ `sedd` requires a special `config.json` file in the current working directory. There's a template available [on GitHub](https://github.com/LunarWatcher/se-data-dump-transformer/blob/master/config.example.json).
57
+
58
+ The only two fields you _need_ to fill out in the template is the email and password fields with credentials for a Stack Exchange account. You need to be logged in to download the data dumps, so the downloader needs the credentials to log in on your behalf. It doesn't matter if you're logged into SE elsewhere, as Selenium automatically creates a blank profile every time it starts, which won't include any cookies from SE, which means login is required.
59
+
60
+ > [!tip]
61
+ >
62
+ > The downloader can automatically create new accounts in the network for you, if you don't have all 180-whatever accounts on every site in the network already. You can also create these by hand if you prefer for some reason, but you are not required to have all 180+ accounts before using the downloader.
63
+
64
+ ## System requirements and pitfalls
65
+
66
+ `sedd` is exclusively Firefox-based, due to Chromium completely gutting support for uBlock Origin and custom filters. You need Firefox installed on your system to use `sedd`.
67
+
68
+ > [!note]
69
+ > On Linux and Windows-based systems, geckodriver is [slightly modified](https://pypi.org/project/undetected-geckodriver-lw/). This is an anti-anti-bot measure meant to prevent Cloudflare loops. If you're on macOS and get sent in a captcha loop, it's recommended you switch to Windows or Linux - a Linux VM is also an option if you have no way out of Apple's closed-down ecosystem.
70
+
71
+ Note that Ubuntu users, or other people who (for whatever reason) choose to use the Snap version of Firefox, have to jump through some extra hoops. The native version of Firefox is strongly encouraged, but if you run into problems with the snap version of Firefox and can't or won't switch, you need to define `export SE_GECKODRIVER=/snap/bin/geckodriver`. Selenium can and will find the snap version of `geckodriver` on its own, but for reasons I simply don't understand, it will still fail with several arbitrary errors.
72
+
73
+ ### Cloudflare issues or download issues.
74
+
75
+ Stack Exchange has configured Cloudflare to be _highly_ aggressive, especially to certain countries. You will almost certainly run into captchas, and the downloader is designed to deal with this. After an initial attempt to solve the captcha on its own, you'll be notified (provided you don't disable the notification provider in `config.json`) and asked to solve it manually.
76
+
77
+ If, at this point, it appears to succeed, but you're redirected back to a full-screen Cloudflare captcha wall, you've likely run into a Cloudflare loop. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cloudflare-loops) for further help. If this doesn't help, please open an issue.
78
+
79
+ If the downloads start fine, but later suddenly fail for no good reason, you're likely running into general download instability. This especially applies to `stackoverflow.com.7z`, as its massive size simply increases the chance you wait for it long enough that it flakes out. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#download-instability-particularly-of-stackoverflowcom7z) for further help.
80
+
81
+ The "Warnings" section in the README may contain additional information about other failure modes not listed here in the future.
82
+
83
+ ## Using the downloader
84
+
85
+ With `./config.json` in the current working directory and Firefox installed, you can now run the downloader with:
86
+ ```python3
87
+ sedd
88
+ ```
89
+
90
+ For command line flags, see `sedd --help`, or [the main readme](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cli-options).
91
+
@@ -0,0 +1,57 @@
1
+ # SE Data Dump Downloader
2
+
3
+
4
+ For more comprehensive information, please read the [main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master) on GitHub. This README contains an abridged version of the main README specifically aimed at Pypi users.
5
+
6
+ For usage problems not listed in this readme, see the main README. If no information exists, please open an issue on GitHub - keeping the tool accessible to everyone is a priority.
7
+
8
+ ---
9
+
10
+ The SE Data Dump Downloader (abbreviated `sedd`) is a command line Selenium-based utility for downloading the entire Stack Exchange data dump in their new [anti-community format](https://stackoverflow.com/help/data-dumps), since they decided not to bother providing an official "download all" button. It's one of two components that operate on the data dump in the second project, the other being the (non-python-based) SE data dump transformer - a project that converts the data dump from the not-so-useful official `.xml` format to some other formats. The pypi package is exclusively for the downloader, and does not ship with a copy of the transformer. See the main README if you're looking for the transformer.
11
+
12
+
13
+ For the pypi version, you can download it with:
14
+ ```python3
15
+ pip3 install sedd
16
+ ```
17
+
18
+ Note that there are some additional steps before you can start using it, that are detailed in this README.
19
+
20
+ ## Configuration
21
+
22
+ `sedd` requires a special `config.json` file in the current working directory. There's a template available [on GitHub](https://github.com/LunarWatcher/se-data-dump-transformer/blob/master/config.example.json).
23
+
24
+ The only two fields you _need_ to fill out in the template is the email and password fields with credentials for a Stack Exchange account. You need to be logged in to download the data dumps, so the downloader needs the credentials to log in on your behalf. It doesn't matter if you're logged into SE elsewhere, as Selenium automatically creates a blank profile every time it starts, which won't include any cookies from SE, which means login is required.
25
+
26
+ > [!tip]
27
+ >
28
+ > The downloader can automatically create new accounts in the network for you, if you don't have all 180-whatever accounts on every site in the network already. You can also create these by hand if you prefer for some reason, but you are not required to have all 180+ accounts before using the downloader.
29
+
30
+ ## System requirements and pitfalls
31
+
32
+ `sedd` is exclusively Firefox-based, due to Chromium completely gutting support for uBlock Origin and custom filters. You need Firefox installed on your system to use `sedd`.
33
+
34
+ > [!note]
35
+ > On Linux and Windows-based systems, geckodriver is [slightly modified](https://pypi.org/project/undetected-geckodriver-lw/). This is an anti-anti-bot measure meant to prevent Cloudflare loops. If you're on macOS and get sent in a captcha loop, it's recommended you switch to Windows or Linux - a Linux VM is also an option if you have no way out of Apple's closed-down ecosystem.
36
+
37
+ Note that Ubuntu users, or other people who (for whatever reason) choose to use the Snap version of Firefox, have to jump through some extra hoops. The native version of Firefox is strongly encouraged, but if you run into problems with the snap version of Firefox and can't or won't switch, you need to define `export SE_GECKODRIVER=/snap/bin/geckodriver`. Selenium can and will find the snap version of `geckodriver` on its own, but for reasons I simply don't understand, it will still fail with several arbitrary errors.
38
+
39
+ ### Cloudflare issues or download issues.
40
+
41
+ Stack Exchange has configured Cloudflare to be _highly_ aggressive, especially to certain countries. You will almost certainly run into captchas, and the downloader is designed to deal with this. After an initial attempt to solve the captcha on its own, you'll be notified (provided you don't disable the notification provider in `config.json`) and asked to solve it manually.
42
+
43
+ If, at this point, it appears to succeed, but you're redirected back to a full-screen Cloudflare captcha wall, you've likely run into a Cloudflare loop. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cloudflare-loops) for further help. If this doesn't help, please open an issue.
44
+
45
+ If the downloads start fine, but later suddenly fail for no good reason, you're likely running into general download instability. This especially applies to `stackoverflow.com.7z`, as its massive size simply increases the chance you wait for it long enough that it flakes out. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#download-instability-particularly-of-stackoverflowcom7z) for further help.
46
+
47
+ The "Warnings" section in the README may contain additional information about other failure modes not listed here in the future.
48
+
49
+ ## Using the downloader
50
+
51
+ With `./config.json` in the current working directory and Firefox installed, you can now run the downloader with:
52
+ ```python3
53
+ sedd
54
+ ```
55
+
56
+ For command line flags, see `sedd --help`, or [the main readme](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cli-options).
57
+
sedd-0.0.0/README.md ADDED
@@ -0,0 +1,296 @@
1
+ # Stack Exchange data dump downloader and transformer
2
+
3
+ [![Data dump transformer build](https://github.com/LunarWatcher/se-data-dump-transformer/actions/workflows/transformer.yml/badge.svg)](https://github.com/LunarWatcher/se-data-dump-transformer/actions/workflows/transformer.yml) [![Stackapps listing](https://img.shields.io/badge/StackApps%20listing-FF9900)](https://stackapps.com/q/10591/69829)
4
+
5
+ **Disclaimer:** This project is not affiliated with Stack Exchange, Inc.
6
+
7
+ ## Background
8
+
9
+ This section contains background on why this project exists. If you know and/or don't care, feel free to skip to the next section.
10
+
11
+ In June 2023, Stack Exchange [briefly cancelled the data dump](https://meta.stackexchange.com/q/389922/332043), and backpedalled after a significant amount of backlash from the community. The status-quo of uploads to archive.org was restored. In the slightly more than a year between June 2023 and July 2024, it looked like they were staying off that path. Notably, they made a [logical shift to the _exact_ dates involved in the upload](https://meta.stackexchange.com/q/398279/) to deal with archive.org being slow. In December 2023, they [announced a delay in the upload](https://meta.stackexchange.com/q/395197/) likely to avoid speculation that another cancellation was happening.
12
+
13
+ We appeared to be out of the woods. But this repo wouldn't exist if that was the case now, would it?
14
+
15
+ ### 2024 data dump restriction attempt
16
+
17
+ In July 2024, Stack Exchange announced the first restrictions on the data dump, by moving it in-house and actively discouraging archive.org reuploads, [likely in violation of the CC-By-SA license](https://meta.stackexchange.com/a/401326/332043)
18
+
19
+ The current revision can be read [here](https://meta.stackexchange.com/q/401324/332043)
20
+
21
+ Here's what's happening:
22
+
23
+ - SE is moving the data dump from archive.org to their own infrastructure
24
+ - They're discontinuing the archive.org dump, which makes it significantly harder to archive the data if SE, for example, were to go out of business
25
+ - In addition to discontinuing the archive.org dump, they're imposing significant restrictions on the data dump
26
+ - They're doing the first revision **without the possibility to download the entire data dump** in one click, a drastic QOL reduction from the current situation.
27
+ - Later in 2024 and in at least the first half of 2025, they've continued scaling up **unreasonably aggressive anti-bot measures**, that also affect the data dump downloader when used in certain countries - even when the browser has full human oversight, and the captchas are solved by a human.
28
+ - At the initial time of the release in 2024, Stack Exchange promised to give download links to the entire data dump to anyone who requests it. A network moderator requested it immediately after release. As of May 22nd, 2025, they're _still_ waiting for SE to actually offer the download. This is in spite of several reminders and SE employees claiming they're going to get on it, and upwards of weekly reminders over a period of at least 4-5 months, after which I personally lost track of the frequency as I resigned as moderator on Stack Overflow. **SE has no interest in providing bulk downloads to the community, regardless of the legitimacy of the request**, and the future of data dump archival rests purely on the community's ability to make tooling that combats SE's ridiculous decisions.
29
+
30
+ This is an opinionated summary of the reason why; SE wants to capitalise on AI companies that need training data, and have decided that the community doesn't matter in that process. The current process, while not nearly as restricting as rev. 1 and 2, is a symptom of precisely one thing; Stack Exchange doesn't care about its users, but rather cares about finding new ways to profit off user data.
31
+
32
+ **Stack Exchange, Inc. is now the single biggest threat to the community**, and to the platform's user-generated and [permissively-licensed content](https://stackoverflow.com/help/licensing) that the community has spent countless hours creating precisely _because_ the data is public.
33
+
34
+ That is why this project exists; this is meant to automate the data dump download process for non-commercial license-compliant use, since Stack Exchange, Inc. couldn't be bothered adding a "download all" button from day 1.
35
+
36
+ As an added bonus, since this project already exists, there's an accompanying system to automatically convert the data dump to other formats. In my experience, the vast majority of applications building on the data dump do not work directly with the XML. Other, more convenient data formats are often created as an intermediate. Aside using it as an intermediate for various forms of analysis, there are a couple major examples of other distribution forms that are listed later in this README.
37
+
38
+ While these are preprocessed distributions of the data dump, this project is also meant to help converting to these various formats. While unlikely to replace the source code for either of these two examples, I hope the transformer system here can get rid of boilerplate for other projects.
39
+
40
+ ## Known archives of new data dumps
41
+
42
+ ### Complete data dump archives
43
+
44
+ A [different project](https://communitydatadump.com/index.html) is currently maintaining a list of both the source data dumps (XML), as well as other distributions. It includes both historical versions of the data dump, as well as new versions uploaded under the new anti-community scheme. In addition, the later data dumps (though exclusively the SE-uploaded ones; not generated variants) [have been uploaded to Academic Torrents](https://academictorrents.com/collection/stack-exchange-data-dumps), which serves as a secondary list.
45
+
46
+ Note that since someone is uploading an unofficial version to archive.org, you may not need to use the downloader at all. However, to make sure this access continues, I strongly encourage you to download directly from SE anyway if you can -- this helps decrease the chance the uploader is identified and blocked by SE, which will turn into a problem for archival efforts in the long term. It may also decrease the chances SE points to low usage numbers as an excuse to axe the data dump entirely.[^4]
47
+
48
+ [^4]: There's no guarantee the data dump will continue existing anymore - removing as many justifications to axe the data dump as possible may become increasingly important at some point. Unfortunately, if it is, we won't find out until it's too late by seeing the data dump get axed.
49
+
50
+ ### Other tools
51
+
52
+ This list contains converter tools that work on all sites and all tables.
53
+
54
+ | Maintainer | Format(s) | First-party torrent available | Converter |
55
+ | ---------- | ----------------------- | ----------------------------- | -------------------------------------------------------------------- |
56
+ | Maxwell175 | SQLite, Postgres, MSSQL | Partially[^2] | [AGPL-3.0](https://github.com/Maxwell175/StackExchangeDumpConverter) |
57
+
58
+ ### Other data dump distributions and conversion tools
59
+
60
+ For completeness (well, sort of, none of these lists are exhaustive), this is a list of incomplete archives (archives that limit the number of included tables and/or sites)
61
+
62
+ | Maintainer | Format | Torrent available | Converter | Site(s) | Tables |
63
+ | ------------ | -------------------------------------------------------------------------------------------------------------- | ----------------- | ------------------------------------------------------ | ------------------- | ---------- |
64
+ | Brent Ozar | [MSSQL](https://www.brentozar.com/archive/2015/10/how-to-download-the-stack-overflow-database-via-bittorrent/) | Yes | [MIT-licensed](https://github.com/BrentOzarULTD/soddi) | Stack Overflow only | All tables |
65
+ | Jason Punyon | [SQLite](https://seqlite.puny.engineering/) | No | Closed-source[^1] | All sites | Posts only |
66
+
67
+ ## Using the downloader
68
+
69
+ Note that it's stongly encouraged that you use a venv. To set one up, run `python3 -m venv env`. After that, you'll need to activate it with one of the activation scripts. Run the appropriate one for your operating system. If you're not sure what the scripts are called, you can find them in `./env/bin`
70
+
71
+ ### Warnings
72
+
73
+ #### Cloudflare loops
74
+
75
+ > [!NOTE]
76
+ > You should be running into this less as long as you don't provide the `-g`/`--disable-undetected-geckodriver` flag. See "Captchas and other misc. barriers" for details. If you haven't disabled it (i.e. you haven't specified either of the two flags) and still run into this problem, continue reading this section.
77
+ >
78
+ > If you're a macOS user, leaving undetected-geckodriver enabled will not help. Switch to Windows or Linux (physically or in a VM) and try again there.
79
+
80
+ ---
81
+
82
+ > TL;DR: If you get a full-page Cloudflare block that loops back to itself after completing the capcha, connect to a VPN, or download the data dump from unofficial community reuploads
83
+
84
+ If you get a full-page Cloudflare block, and solving the captcha redirects you right back to the cloudflar eblock page even if you complete it correctly, you have to switch to a VPN in another country. For reasons beyond me, SE hard-blocks automated browsers _from specific countries only_. One of the verified blocked countries used for this test was Singapore.
85
+
86
+ If you get slapped with a Cloudflare loop, the only option for now is to use a VPN in another country. Switzerland and Norway have both been verified to work at the time of writing. Fascinatingly, using a VPN makes no difference on the looping; it's purely country-based, not anti-VPN-based. The loop has been verified on both a residential IP and a datacenter IP (VPN).
87
+
88
+ I unfortunately do not (and cannot) write a complete list of countries affected by this bullshit, so you have to test this manually. If you do not have access to a VPN, check https://communitydatadump.com/ or https://academictorrents.com/collection/stack-exchange-data-dumps for archived (unofficial) versions uploaded by the community. They're usually uploaded within a few days to a couple weeks, and unless you're downloading directly from the archive.org version, is significantly faster and more stable than downloading from SE themselves.
89
+
90
+ #### Download instability, particularly of `stackoverflow.com.7z`
91
+
92
+ > TL;DR: If `stackoverflow.com.7z` fails to download:
93
+ > 1. You're on a good connection: consider connecting to a VPN before retrying.
94
+ > 2. You're on a bad connection: consider downloading via the unofficial community torrent instead. It's significantly more resistant to problems caused by bad internet connections.
95
+
96
+ If you keep running into the download of `stackoverflow.com.7z` failing, you may need to connect to a VPN. For reasons that are not clear (but that are likely down to SE being horrible at implementing things that work), `stackoverflow.com.7z` can just randomly fail. This assumes your device has a stable internet connection for the duration of the download, but arbitrarily fails anyway.
97
+
98
+ For some reason, connecting with a VPN offers just enough extra stability to avoid whatever causes the failures. This is especially the case if you're on a relatively slow (100Mbps) network, as longer download times increases the chance of failure. If you're on a generally unstable network, a VPN may help, but if it doesn't, it will be significantly easier to download the unofficial community reuploads, as these are always available as torrents. Torrents are generally more resistant to full failure on bad network connections, and can download single missing pieces rather than forcing you to redownload the entire 68+GB thing.
99
+
100
+ ### Requirements
101
+
102
+ - Python 3.10 or newer[^3]
103
+ - `pip3 install -r requirements.txt`
104
+ - Lots of storage. The 2024Q1 data dump was 92GB compressed.
105
+ - A display you can access somehow (physical or virtual, but you need to be able to see it) to be able to solve captchas
106
+ - Email and password login for Stack Exchange - Google, Facebook, GitHub, and other login methods are not supported, and will not be supported.
107
+ - If you don't have this, see [this meta question](https://meta.stackexchange.com/a/1847/332043) for instructions.
108
+ - Firefox installed
109
+ - Snap and flatpak users may run into problems; it's strongly recommended to have a non-snap/flatpak installation of Firefox and Geckodriver.
110
+ - Known errors:
111
+ - "The geckodriver version may not be compatible with the detected firefox version" - update Firefox and Geckodriver. If this still doesn't work, consider switching to a non-snap installation of Firefox and Geckodriver.
112
+ - "Your Firefox profile cannot be loaded" - One of Geckodriver or Firefox is Snap-based, while the other is not. [Consider switching to a non-snap installation](https://stackoverflow.com/a/72531719/6296561) of Firefox, or verifying that your PATH is set correctly.
113
+ - If you need to manaully install Geckodriver (which shouldn't normally be necessary; it's often bundled with Firefox in one way or another), the binaries are on [GitHub](https://github.com/mozilla/geckodriver/releases)
114
+
115
+ The downloader does **not** support Docker due to the display requirement.
116
+
117
+ ### Config, running, and what to expect
118
+
119
+ #### Configuring and starting
120
+
121
+ 1. Make sure you have all the requirements from the Requirements section.
122
+ 2. Copy `config.example.json` to `config.json`
123
+ 3. Open `config.json`, and edit in the values.
124
+ 4. Run the extractor with `python3 -m sedd`. If you're on Windows, you may need to run `python -m sedd` instead.
125
+
126
+ > [!TIP]
127
+ >
128
+ > If you haven't recently updated, it's recommended that you `git pull && pip3 install -r requirements.txt` first. This is not required if this is the first time you're running the tool, as everything will be up-to-date. This goes from being optional to a very good idea or outright required if you run into hard Cloudflare blocks you can't get around.
129
+
130
+ ##### Config option values
131
+
132
+ ###### `notifications.provider`
133
+ Supported values:
134
+ * `"native"`: Uses your operating system's native notification method
135
+ * `null`: Disables notifications
136
+
137
+ #### CLI options
138
+
139
+ Exractor CLI supports the following configuration options:
140
+
141
+ | Short | Long | Type | Default | Description |
142
+ | ----- | ---------------------- | -------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
143
+ | `-o` | `--outputDir <path>` | Optional | `<cwd>/downloads` | Specifies the directory to download the archives to. |
144
+ | `-k` | `--keep-consent` | Optional | `false` | Whether to keep OneTrust's consent dialog. If set, you are responsible for getting rid of it yourself (uBlock can handle that for you too). |
145
+ | `-s` | `--skip-loaded <path>` | Optional | - | Whether to skip over archives that have already been downloaded. An archive is considered to be downloaded if the output directory has one already & the file is not empty. |
146
+ | - | `--dry-run` | Optional | - | Whether to actually download the archives. If set, only traverses the network's sites. |
147
+ | `-g` | `--disable-undetected-geckodriver` | Optional | `false` | Whether or not to disable the anti-anti-bot detection webdriver, and use a fully standard Firefox instead. See "Captchas and other misc. barriers" for more information. Has no effect on macOS. |
148
+ | `-v` | - | Optional | `false` | Whether or not to enable verbose logging. You do not want to enable this unless you're diagnosing a problem with sedd - the verbose logging includes low-level Selenium output. |
149
+
150
+ #### Captchas and other misc. barriers
151
+
152
+ This software is designed around Selenium, a browser automation tool. This does, however, mean that the program can be stopped by various bot defenses. This would happen even if you downloaded all the [~183 data dumps](https://stackexchange.com/sites#questionsperday) fully by hand, because it's a _lot_ of repeated operations.
153
+
154
+ This is where notification systems come in; expecting you to sit and watch for potentially a significant number of hours is not a good use of time. If anything happens, you'll be notified, so you don't have to continuously watch the program. Currently, only a native desktop notifier is supported, but support for other notifiers may be added in the future.
155
+
156
+ Unfortunately, since the initial release of sedd, Stack Exchange, Inc. has bumped their bot blocking to outright unreasonable levels, which has affected sedd from certain countries. As of 2.0.0, non-macOS users automatically use a non-standard version of Firefox. The only modification done is that the `navigator.webdriver` property is disabled, which is enough to downgrade infinite captcha loops to a single captcha, or no blocking at all.
157
+
158
+ This can be disabled via the `-g`/`--disable-undetected-geckodriver` flag, **but leaving it on is strongly encouraged**. The alternate webdriver used disables one of the obvious browser-level bot detection flags that Cloudflare uses to send the browser in an infinite captcha loop. Disabling this, depending on how aggressive Stack Exchange's bot blocking has got since this was written, may mean you cannot download the data dump without a VPN into another country (see [Cloudflare loops](#cloudflare-loops))
159
+
160
+ <details>
161
+ <summary>
162
+ Technical explanation (and a stab at Stack Exchange, Inc.)
163
+ </summary>
164
+ As of 2.0.0, <a href="https://github.com/LunarWatcher/undetected_geckodriver/">undetected-geckodriver-lw</a> is used. As described, it's functionally identical to the standard Selenium Firefox driver, except it disables <code>navigator.webdriver</code>. In the future, additional anti-anti-bot measures may be implemented as well, depending on how much needs to be done low-level to avoid detection.
165
+
166
+ All this does, however, is identify the browser as most likely human. It can and still will run into captcha walls whenever Cloudflare feels like it. But all the obvious, hard, self-announced "hey, I'm a bot" identifiers in the browser are now gone. This is a full nuclear response to SE's nuclear anti-bot actions, that also took down the data dump downloader in several major countries.
167
+
168
+ Your move, Stack Exchange. I have several contingencies planned if necessary - you will not kill the data dump downloader in any other way than outright killing the data dump itself.
169
+ </details>
170
+
171
+ If you don't run into Cloudflare loops, but instead run into hard blocks or other rate limiting, please [open an issue](https://github.com/LunarWatcher/undetected_geckodriver/issues). These blocks are a threat to archival, and I will happily continue to fight against them as far as necessary to ensure access to the data dumps.
172
+
173
+ #### Storage space
174
+
175
+ As of Q1 2024, the data dump was a casual 93GB in compressed size. If you have your own system to transform the data dump after downloading, you only need to worry about the raw size of the data dump.
176
+
177
+ However, if you use the built-in transformer pipeline, you'll need to expect a _lot_ more data use. Depending on how much concurrent processing you do, you may need up to 3-700GB of storage space. It's recommended you have around 1TB **free** on disk to avoid running out, especially if you're converting to SQLite. JSON or other formats that don't merge `stackoverflow.7z` into a single output file can probably get away with around 500GB.
178
+
179
+ The output, by default, is compressed back into 7z if dealing with a file-based transformer. Due to this, an intermediate file write is performed prior to compressing back into a .7z. At runtime, you need:
180
+
181
+ - The compressed data dump; at least 92GB and increasing with each dump
182
+ - The compressed converted data dump; depending on compression rates for the specific format, this anywhere from a little less than the original size to significantly larger
183
+ - A significant amount of space for intermediate files. While these will be deleted as soon as they're done and compressed, they'll take up a significant amount of space on the disk in the meanwhile
184
+
185
+ Note that the transformer pipeline is executed separately; see the transformer section below.
186
+
187
+ No other formats than .7z are supported at this time, but are planned for some point in the future.
188
+
189
+ #### Execution time
190
+
191
+ One of the major downsides with the way this project functions is that it's subject to Cloudflare bullshit. This means that the total time to download is `(combined size of data dumps) / (internet speed) + (rate limiting) + (navigation overhead) + (time to solve captchas)`. While navigation overhead and rate limiting (hopefully) doesn't account for a significant share of time, it can potentially be significant. It's certainly a slower option than archive.org's torrent.
192
+
193
+
194
+ ## Using the transformer
195
+
196
+ Once you've downloaded the data dumps, you may want to transform it into a more usable format than the data dump offers by default. This is where the transformer component comes in.
197
+
198
+ ### Docker
199
+
200
+ This section assumes you have Docker installed, with [docker-compose-v2](https://docs.docker.com/compose/migrate/).
201
+
202
+ A [default compose file](./docker-compose.yml) is provided for convenience. If you want to use a different one, set the `COMPOSE_FILE` environment variable in a custom `.env` file in the project's root directory, or provide it on the command line. To start the container (and build it if not done already), run:
203
+
204
+ ```bash
205
+ docker compose up
206
+ ```
207
+
208
+ By default, this binds `downloads` and `out` in the current working directory to the container. If you want to change these paths, update the `volumes` attribute mapping of the `transformer` service in your compose file (or the default docker-compose.yml file if not using a custom one).
209
+
210
+ #### Environment variables
211
+
212
+ The following environment variables are used:
213
+
214
+ | Name | Values | Default | Description |
215
+ | ------------------ | -------------- | ------- | ------------------------- |
216
+ | `SEDD_OUTPUT_TYPE` | `json\|sqlite` | `json` | Any upported output type. |
217
+ | `SPDLOG_LEVEL` | log level | `info` | Sets the logging level. |
218
+
219
+ If you want to override the defaults, either set them in a custom `.env` file in the project's root directory, or, if you have a UNIX shell (i.e. not cmd or PowerShell; Windows users can use Git Bash), you can explicitly provide them on the command line:
220
+
221
+ ```bash
222
+ SEDD_OUTPUT_TYPE=sqlite docker compose up
223
+ ```
224
+
225
+ If you want to rebuild the container, pass the `--build` flag to the docker command.
226
+
227
+ If you insist on using cmd or PowerShell instead of a good shell, setting the variables is left as an exercise to the reader.
228
+
229
+ ### Native
230
+
231
+ #### Requirements
232
+
233
+ - C++20 compiler
234
+ - CMake 3.10 or newer
235
+
236
+ Other dependencies (stc, libarchive, spdlog, and pugixml) are automatically handled by CMake using FetchContent. Unlike the downloader, this component can run without a display.
237
+
238
+ #### Running
239
+
240
+ TL;DR:
241
+
242
+ ```bash
243
+ cd transformer
244
+ mkdir build
245
+ cd build
246
+ # Option 1: debug:
247
+ cmake .. -DCMAKE_BUILD_TYPE=Debug
248
+ # Option 2: release mode; strongly recommended for anything that needs the performance:
249
+ cmake .. -DCMAKE_BUILD_TYPE=Release
250
+ # ---
251
+ # Replace 8 with the number of cores/threads you have
252
+ cmake --build . -j 8
253
+
254
+ # Note: this only works after running the Python downloader
255
+ # For early testing, I've been populating this folder with
256
+ # files from the old archive.org data dump.
257
+ # The last argument is the path to the downloaded data
258
+ # *UNIX:
259
+ ./sedd-transformer -i ../../downloads -t [formatter type]
260
+ # Windows
261
+ .\sedd-transformer.exe -i ..\..\downloads -t [formatter type]
262
+ ```
263
+
264
+ Pass `--help` to see the available formatters for your current version of the data dump transformer.
265
+
266
+ ### Supported transformers
267
+
268
+ Currently, the following transformers are supported:
269
+
270
+ - `json`
271
+ - `sqlite`
272
+ - Note: All data related to a site is merged into a single database
273
+
274
+ ## Language rationale
275
+
276
+ While I really didn't want to split the system over two programming languages, this is unfortunately the best way to go about it.
277
+
278
+ C++ does not really support Selenium, which is effectively a requirement for the download process. There are bindings, but all of them appear to be out-of-date, and I don't feel like writing an entire system for selenium
279
+
280
+ Python, on the other hand, infuriatingly doesn't support 7z streaming, at least not in a convenient format. There's the `libarchive` package, but it refuses to build. `python-libarchive` allegedly does, but [Windows support is flaky](https://github.com/smartfile/python-libarchive/issues/38), so the transformer might've had to be separated from the downloader anyway. There's py7zr, which does work everywhere, but it [doesn't support 7z streaming](https://github.com/miurahr/py7zr/issues/579).
281
+
282
+ 7z and XML streaming are both _critical_ for the processing pipeline. If you plan to convert the entire data dump, you'll eventually run into `stackoverflow.com-PostHistory.7z`, which is 39GB compressed, and **181GB uncompressed** in the 2024 Q1 data dump. As time passes, this will likely continue to grow, and the absurd amounts of RAM required to just tank the full size [is barely supported on modern and _very_ high-end hardware](https://www.reddit.com/r/buildapc/comments/17hqk3k/what_happened_to_256gb_ram_capacity_motherboards/). Finding someone able to tank that is going to be difficult for the vast majority of people.
283
+
284
+ Consequently, direct `libarchive` support is beneficial, and rather than writing an entire new python wrapper (or taking over an existing one), it's easier to just write that part in C++. Also, since it might be easier to run this particular part in a Docker container to avoid downloading build tools on certain systems, having it be fully headless is an advantage.
285
+
286
+ On the bright side, this should mean faster processing compared to Python.
287
+
288
+ ## License
289
+
290
+ The code is under the MIT license; see the `LICENSE` file.
291
+
292
+ The data downloaded and produced is under various versions of [CC-By-SA](https://stackoverflow.com/help/licensing), as per Stack Exchange's licensing rules, in addition to whatever extra rules they try to impose on the data dump.
293
+
294
+ [^1]: I've been unable to find the generator code, but I've also been unable to find a statement confirming that it's closed-source. It's possible it is open-source, but if it is, it's hard to find the source
295
+ [^2]: Only Postgres at the time of writing, with more planned
296
+ [^3]: Might work with earlier versions, but these are untested and not supported
@@ -0,0 +1,56 @@
1
+ [build-system]
2
+ requires = ["setuptools >= 77.0.3"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [tool.setuptools]
6
+ packages = ["sedd"]
7
+
8
+ [tool.setuptools-git-versioning]
9
+ enabled = true
10
+
11
+ [project]
12
+ name = "sedd"
13
+ dynamic = ["version"]
14
+ license = "MIT"
15
+ license-files = ["LICENSE"]
16
+ dependencies = [
17
+ "selenium==4.32.0",
18
+ "undetected-geckodriver-lw>=2.1.0",
19
+ "desktop-notifier==5.0.1",
20
+ "watchdog==4.0.2",
21
+ "requests"
22
+ ]
23
+ readme = "README-python.md"
24
+
25
+ authors = [
26
+ { name = "LunarWatcher", email = "oliviawolfie@pm.me" },
27
+ ]
28
+
29
+ description = "Unofficial, community-made tool for downloading the Stack Exchange data dumps"
30
+ requires-python = ">=3.10"
31
+ classifiers = [
32
+ "Development Status :: 5 - Production/Stable",
33
+ "Environment :: Console",
34
+ "Intended Audience :: Developers",
35
+ "Topic :: System :: Archiving",
36
+ "Topic :: Utilities",
37
+ "Programming Language :: Python :: 3",
38
+ "Programming Language :: Python :: 3.10",
39
+ "Programming Language :: Python :: 3.11",
40
+ "Programming Language :: Python :: 3.12",
41
+ "Operating System :: POSIX :: Linux",
42
+ "Operating System :: MacOS",
43
+ "Operating System :: Microsoft :: Windows",
44
+ ]
45
+ keywords = [
46
+ "Stack Exchange", "data dump", "Stack Exchange data dump",
47
+ "downloader"
48
+ ]
49
+
50
+ [project.urls]
51
+ Homepage = "https://github.com/LunarWatcher/se-data-dump-transformer"
52
+ Documentation = "https://github.com/LunarWatcher/se-data-dump-transformer"
53
+ Repository = "https://github.com/LunarWatcher/se-data-dump-transformer.git"
54
+ Issues = "https://github.com/LunarWatcher/se-data-dump-transformer/issues"
55
+ Changelog = "https://github.com/LunarWatcher/se-data-dump-transformer/releases"
56
+
File without changes
@@ -0,0 +1 @@
1
+ from .main import *
sedd-0.0.0/sedd/cli.py ADDED
@@ -0,0 +1,69 @@
1
+ import argparse
2
+
3
+ from os import getcwd
4
+ from os.path import join
5
+
6
+
7
+ class SEDDCLIArgs(argparse.Namespace):
8
+ skip_loaded: bool
9
+ keep_consent: bool
10
+ output_dir: str
11
+ dry_run: bool
12
+ disable_undetected: bool
13
+ verbose: bool
14
+
15
+
16
+ parser = argparse.ArgumentParser(
17
+ prog="sedd",
18
+ description="Automatic (unofficial) SE data dump downloader for the anti-community data dump format",
19
+ )
20
+
21
+ parser.add_argument(
22
+ "-s", "--skip-loaded",
23
+ required=False,
24
+ default=False,
25
+ action="store_true",
26
+ dest="skip_loaded"
27
+ )
28
+
29
+ parser.add_argument(
30
+ "-k", "--keep-consent",
31
+ required=False,
32
+ dest="keep_consent",
33
+ action="store_true",
34
+ default=False
35
+ )
36
+
37
+ parser.add_argument(
38
+ "-o", "--outputDir",
39
+ required=False,
40
+ dest="output_dir",
41
+ default=join(getcwd(), "downloads")
42
+ )
43
+
44
+ parser.add_argument(
45
+ "-g", "--disable-undetected-geckodriver",
46
+ required=False,
47
+ dest="disable_undetected",
48
+ action="store_true",
49
+ default=False
50
+ )
51
+
52
+ parser.add_argument(
53
+ "--dry-run",
54
+ required=False,
55
+ default=False,
56
+ action="store_true",
57
+ dest="dry_run"
58
+ )
59
+ parser.add_argument(
60
+ "-v",
61
+ required=False,
62
+ default=False,
63
+ action="store_true",
64
+ dest="verbose"
65
+ )
66
+
67
+
68
+ def parse_cli_args() -> SEDDCLIArgs:
69
+ return parser.parse_args()
@@ -0,0 +1,62 @@
1
+ from os import path, makedirs
2
+ from urllib import request
3
+ from json import dumps
4
+ from uuid import uuid4
5
+ import platform
6
+
7
+ from selenium import webdriver
8
+ from selenium.webdriver.firefox.options import Options
9
+ from undetected_geckodriver import Firefox as UFirefox
10
+
11
+ from .config import SEDDConfig
12
+ from .ubo import init_ubo_settings
13
+
14
+
15
+ def init_output_dir(output_dir: str):
16
+ if not path.exists(output_dir):
17
+ makedirs(output_dir)
18
+
19
+ print(output_dir)
20
+
21
+ return output_dir
22
+
23
+
24
+ def init_firefox_driver(config: SEDDConfig, disable_undetected: bool, output_dir: str):
25
+ options = Options()
26
+ options.enable_downloads = True
27
+ options.set_preference("browser.download.folderList", 2)
28
+ options.set_preference("browser.download.manager.showWhenStarting", False)
29
+ options.set_preference("browser.download.dir", output_dir)
30
+ options.set_preference(
31
+ "browser.helperApps.neverAsk.saveToDisk", "application/x-gzip"
32
+ )
33
+
34
+ # our own uuid for uBO so as we don't need to do the dance of inspecing internals
35
+ ubo_internal_uuid = f"{uuid4()}"
36
+
37
+ options.set_preference("extensions.webextensions.uuids", dumps(
38
+ {"uBlock0@raymondhill.net": ubo_internal_uuid}))
39
+
40
+ is_apple = platform.system() == "Darwin"
41
+ use_undetected = not disable_undetected and not is_apple
42
+ if use_undetected:
43
+ print("Using undetected-geckodriver")
44
+ browser = UFirefox(options = options)
45
+ else:
46
+ print("Warning: using standard geckodriver. Cloudflare may perpetually block you")
47
+ if is_apple:
48
+ print("This option is forced on macOS. For undetected_geckodriver, "
49
+ "run the downloader in a Linux or Windows environment.")
50
+ browser = webdriver.Firefox(options=options)
51
+
52
+ ubo_download_url = config.get_ubo_download_url()
53
+
54
+ if not path.exists("ubo.xpi"):
55
+ print(f"Downloading uBO from: {ubo_download_url}")
56
+ request.urlretrieve(ubo_download_url, "ubo.xpi")
57
+
58
+ ubo_id = browser.install_addon("ubo.xpi", temporary=True)
59
+
60
+ init_ubo_settings(browser, config, ubo_internal_uuid)
61
+
62
+ return browser, ubo_id
@@ -0,0 +1,263 @@
1
+ from selenium.webdriver.common.by import By
2
+ from selenium.webdriver.firefox.webdriver import WebDriver
3
+ from selenium.common.exceptions import NoSuchElementException
4
+ from typing import Dict
5
+
6
+
7
+ from time import sleep
8
+
9
+ import re
10
+ import sys
11
+ from traceback import print_exception
12
+
13
+
14
+ from .cli import parse_cli_args
15
+ from .config import load_sedd_config
16
+ from .data import sites
17
+ from .meta import notifications
18
+ from .watcher.observer import register_pending_downloads_observer
19
+ from . import utils
20
+
21
+ from .driver import init_output_dir, init_firefox_driver
22
+ import logging
23
+
24
+ args = parse_cli_args()
25
+
26
+ if args.verbose:
27
+ logging.basicConfig(level=logging.DEBUG)
28
+
29
+ sedd_config = load_sedd_config()
30
+
31
+ output_dir = init_output_dir(args.output_dir)
32
+
33
+ browser, ubo_id = init_firefox_driver(
34
+ sedd_config,
35
+ args.disable_undetected,
36
+ output_dir
37
+ )
38
+
39
+
40
+ def kill_cookie_shit(browser: WebDriver):
41
+ sleep(3)
42
+ browser.execute_script(
43
+ """let elem = document.getElementById("onetrust-banner-sdk"); if (elem) { elem.parentNode.removeChild(elem); }""")
44
+ sleep(1)
45
+
46
+ def check_cloudflare_intercept(browser: WebDriver):
47
+ if browser.title == "Just a moment...":
48
+ print("CF verification hit. Trying soft workaround")
49
+ sleep(15)
50
+
51
+ if (browser.title == "Just a moment..."):
52
+ print("Irrecoverable state suspected; captcha solving likely required")
53
+ notifications.notify("CloudFlare verification hit; auto-verification failed. Please complete the captcha", sedd_config)
54
+ else:
55
+ print("Auto-recovered from CF wall")
56
+ return
57
+
58
+ while browser.title == "Just a moment...":
59
+ print("Still stuck on CF verification. Waiting for 10 seconds")
60
+ sleep(10)
61
+
62
+ def is_logged_in(browser: WebDriver, site: str):
63
+ url = f"{site}/users/current"
64
+ browser.get(url)
65
+ sleep(1)
66
+ check_cloudflare_intercept(browser)
67
+
68
+ return "/users/" in browser.current_url
69
+
70
+
71
+ def login_or_create(browser: WebDriver, site: str):
72
+ if is_logged_in(browser, site):
73
+ print("Already logged in")
74
+ else:
75
+ print("Not logged in and/or not registered. Logging in now")
76
+ while True:
77
+ browser.get(f"{site}/users/login")
78
+ check_cloudflare_intercept(browser)
79
+
80
+ if "?newreg" in browser.current_url:
81
+ print(f"Auto-created {site} without login needed")
82
+ break
83
+
84
+ email_elem = browser.find_element(By.ID, "email")
85
+ password_elem = browser.find_element(By.ID, "password")
86
+ email_elem.send_keys(sedd_config.email)
87
+ password_elem.send_keys(sedd_config.password)
88
+ retryLogin = False
89
+
90
+ curr_url = browser.current_url
91
+ browser.find_element(By.ID, "submit-button").click()
92
+
93
+ try:
94
+ elem = browser.find_element(By.CSS_SELECTOR, "#login-form > .js-error-message")
95
+ if elem is not None:
96
+ print("Login failed quietly. Retrying")
97
+ continue
98
+ except:
99
+ # No error element
100
+ pass
101
+
102
+ check_cloudflare_intercept(browser)
103
+
104
+ while browser.current_url == curr_url:
105
+ sleep(3)
106
+
107
+
108
+ captcha_walled = False
109
+ while "/nocaptcha" in browser.current_url:
110
+ if not captcha_walled:
111
+ captcha_walled = True
112
+
113
+ notifications.notify(
114
+ "Captcha wall hit during login", sedd_config
115
+ )
116
+
117
+ sleep(10)
118
+
119
+ if captcha_walled or retryLogin:
120
+ continue
121
+
122
+ if not is_logged_in(browser, site):
123
+ raise RuntimeError("Login failed")
124
+
125
+ break
126
+
127
+
128
+ def download_data_dump(browser: WebDriver, site: str, meta_url: str, etags: Dict[str, str]):
129
+ print(f"Downloading data dump from {site}")
130
+
131
+ def _exec_download(browser: WebDriver):
132
+ if args.keep_consent:
133
+ print('Consent dialog will not be auto-removed')
134
+ else:
135
+ kill_cookie_shit(browser)
136
+
137
+ try:
138
+ checkbox = browser.find_element(By.ID, "datadump-agree-checkbox")
139
+ btn = browser.find_element(By.ID, "datadump-download-button")
140
+ except NoSuchElementException:
141
+ raise RuntimeError(f"Bad site: {site}")
142
+
143
+ if args.dry_run:
144
+ return
145
+
146
+ browser.execute_script("""
147
+ (function() {
148
+ let oldFetch = window.fetch;
149
+ window.fetch = (url, opts) => {
150
+ let promise = oldFetch(url, opts);
151
+
152
+ if (url.includes("/link")) {
153
+ promise.then(res => {
154
+ res.clone().json().then(json => {
155
+ window.extractedUrl = json["url"];
156
+ console.log(extractedUrl);
157
+ });
158
+ return res;
159
+ });
160
+ return new Promise(resolve => setTimeout(resolve, 4000))
161
+ .then(_ => promise);
162
+ }
163
+ return promise;
164
+ };
165
+ })();
166
+ """)
167
+
168
+ checkbox.click()
169
+ sleep(1)
170
+ btn.click()
171
+ check_cloudflare_intercept(browser)
172
+ sleep(2)
173
+ url = browser.execute_script("return window.extractedUrl;")
174
+ utils.extract_etag(url, etags)
175
+
176
+ sleep(5)
177
+
178
+ main_loaded = utils.is_file_downloaded(args.output_dir, site)
179
+ meta_loaded = utils.is_file_downloaded(args.output_dir, meta_url)
180
+
181
+ if not args.skip_loaded or not main_loaded or not meta_loaded:
182
+ if args.skip_loaded and main_loaded:
183
+ pass
184
+ else:
185
+ browser.get(f"{site}/users/data-dump-access/current")
186
+ check_cloudflare_intercept(browser)
187
+
188
+ if not args.dry_run:
189
+ utils.archive_file(args.output_dir, site)
190
+
191
+ _exec_download(browser)
192
+
193
+ if args.skip_loaded and meta_loaded:
194
+ pass
195
+ else:
196
+ browser.get(f"{meta_url}/users/data-dump-access/current")
197
+ check_cloudflare_intercept(browser)
198
+
199
+ if not args.dry_run:
200
+ utils.archive_file(args.output_dir, meta_url)
201
+
202
+ _exec_download(browser)
203
+
204
+
205
+ etags: Dict[str, str] = {}
206
+
207
+ try:
208
+ state, observer = register_pending_downloads_observer(args.output_dir)
209
+
210
+ for site in sites.sites:
211
+ if site not in ["https://meta.stackexchange.com", "https://stackapps.com"]:
212
+ # https://regex101.com/r/kG6nTN/1
213
+ meta_url = re.sub(
214
+ r"(https://(?:[^.]+\.(?=stackexchange))?)", r"\1meta.", site)
215
+
216
+ main_loaded = utils.is_file_downloaded(args.output_dir, site)
217
+ meta_loaded = utils.is_file_downloaded(args.output_dir, meta_url)
218
+
219
+ if args.skip_loaded and main_loaded and meta_loaded:
220
+ pass
221
+ else:
222
+ print(f"Extracting from {site}...")
223
+
224
+ login_or_create(browser, site)
225
+ download_data_dump(
226
+ browser,
227
+ site,
228
+ meta_url,
229
+ etags
230
+ )
231
+
232
+ if observer:
233
+ pending = state.size()
234
+
235
+ print(f"Waiting for {pending} download{'s'[:pending^1]} to complete")
236
+
237
+ while True:
238
+ if state.empty():
239
+ observer.stop()
240
+ browser.quit()
241
+
242
+ utils.cleanup_archive(args.output_dir)
243
+ break
244
+ else:
245
+ sleep(1)
246
+
247
+ except KeyboardInterrupt:
248
+ pass
249
+
250
+ except:
251
+ exception = sys.exc_info()
252
+
253
+ try:
254
+ print_exception(exception)
255
+ except:
256
+ print(exception)
257
+
258
+ browser.quit()
259
+ finally:
260
+ # TODO: replace with validation once downloading is verified done
261
+ # (or export for separate, later verification)
262
+ # Though keeping it here, removing files and re-running downloads feels like a better idea
263
+ print(etags)
@@ -0,0 +1,85 @@
1
+ from typing import Dict
2
+ import requests as r
3
+ from urllib.parse import urlparse
4
+ import os.path
5
+ import re
6
+ import sys
7
+
8
+ from .data.files_map import files_map, inverse_files_map
9
+ from .data.sites import sites
10
+
11
+
12
+ def extract_etag(url: str, etags: Dict[str, str]):
13
+ res = r.get(
14
+ url,
15
+ stream=True
16
+ )
17
+ if res.status_code != 200:
18
+ raise RuntimeError(f"Panic: failed to get {url}: {res.status_code}")
19
+
20
+ etag = res.headers["ETag"]
21
+ res.close()
22
+
23
+ parsed_url = urlparse(url)
24
+ path = parsed_url.path
25
+ filename = os.path.basename(path)
26
+
27
+ etags[filename] = etag
28
+
29
+ print(f"ETag for {filename}: {etag}")
30
+
31
+
32
+ def get_file_name(site_or_url: str) -> str:
33
+ domain = re.sub(r'https://', '', site_or_url)
34
+
35
+ try:
36
+ file_name = files_map[domain]
37
+ return f'{file_name}.7z'
38
+ except KeyError:
39
+ return f'{domain}.7z'
40
+
41
+
42
+ def is_dump_file(file_name: str) -> bool:
43
+ file_name = re.sub(r'\.7z$', '', file_name)
44
+
45
+ try:
46
+ inverse_files_map[file_name]
47
+ except KeyError:
48
+ origin = f'https://{file_name}'
49
+ return origin in sites
50
+
51
+ return True
52
+
53
+
54
+ def check_file(base_path: str, file_name: str) -> bool:
55
+ try:
56
+ res = os.stat(os.path.join(base_path, file_name))
57
+ return res.st_size > 0
58
+ except FileNotFoundError:
59
+ return False
60
+
61
+
62
+ def archive_file(base_path: str, site_or_url: str) -> None:
63
+ try:
64
+ file_name = get_file_name(site_or_url)
65
+ file_path = os.path.join(base_path, file_name)
66
+ os.rename(file_path, f"{file_path}.old")
67
+ except FileNotFoundError:
68
+ pass
69
+
70
+
71
+ def cleanup_archive(base_path: str) -> None:
72
+ try:
73
+ file_entries = os.listdir(base_path)
74
+
75
+ for entry in file_entries:
76
+ if entry.endswith('.old'):
77
+ entry_path = os.path.join(base_path, entry)
78
+ os.remove(entry_path)
79
+ except:
80
+ print(sys.exc_info())
81
+
82
+
83
+ def is_file_downloaded(base_path: str, site_or_url: str) -> bool:
84
+ file_name = get_file_name(site_or_url)
85
+ return check_file(base_path, file_name)
@@ -0,0 +1,91 @@
1
+ Metadata-Version: 2.4
2
+ Name: sedd
3
+ Version: 0.0.0
4
+ Summary: Unofficial, community-made tool for downloading the Stack Exchange data dumps
5
+ Author-email: LunarWatcher <oliviawolfie@pm.me>
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/LunarWatcher/se-data-dump-transformer
8
+ Project-URL: Documentation, https://github.com/LunarWatcher/se-data-dump-transformer
9
+ Project-URL: Repository, https://github.com/LunarWatcher/se-data-dump-transformer.git
10
+ Project-URL: Issues, https://github.com/LunarWatcher/se-data-dump-transformer/issues
11
+ Project-URL: Changelog, https://github.com/LunarWatcher/se-data-dump-transformer/releases
12
+ Keywords: Stack Exchange,data dump,Stack Exchange data dump,downloader
13
+ Classifier: Development Status :: 5 - Production/Stable
14
+ Classifier: Environment :: Console
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Topic :: System :: Archiving
17
+ Classifier: Topic :: Utilities
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3.10
20
+ Classifier: Programming Language :: Python :: 3.11
21
+ Classifier: Programming Language :: Python :: 3.12
22
+ Classifier: Operating System :: POSIX :: Linux
23
+ Classifier: Operating System :: MacOS
24
+ Classifier: Operating System :: Microsoft :: Windows
25
+ Requires-Python: >=3.10
26
+ Description-Content-Type: text/markdown
27
+ License-File: LICENSE
28
+ Requires-Dist: selenium==4.32.0
29
+ Requires-Dist: undetected-geckodriver-lw>=2.1.0
30
+ Requires-Dist: desktop-notifier==5.0.1
31
+ Requires-Dist: watchdog==4.0.2
32
+ Requires-Dist: requests
33
+ Dynamic: license-file
34
+
35
+ # SE Data Dump Downloader
36
+
37
+
38
+ For more comprehensive information, please read the [main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master) on GitHub. This README contains an abridged version of the main README specifically aimed at Pypi users.
39
+
40
+ For usage problems not listed in this readme, see the main README. If no information exists, please open an issue on GitHub - keeping the tool accessible to everyone is a priority.
41
+
42
+ ---
43
+
44
+ The SE Data Dump Downloader (abbreviated `sedd`) is a command line Selenium-based utility for downloading the entire Stack Exchange data dump in their new [anti-community format](https://stackoverflow.com/help/data-dumps), since they decided not to bother providing an official "download all" button. It's one of two components that operate on the data dump in the second project, the other being the (non-python-based) SE data dump transformer - a project that converts the data dump from the not-so-useful official `.xml` format to some other formats. The pypi package is exclusively for the downloader, and does not ship with a copy of the transformer. See the main README if you're looking for the transformer.
45
+
46
+
47
+ For the pypi version, you can download it with:
48
+ ```python3
49
+ pip3 install sedd
50
+ ```
51
+
52
+ Note that there are some additional steps before you can start using it, that are detailed in this README.
53
+
54
+ ## Configuration
55
+
56
+ `sedd` requires a special `config.json` file in the current working directory. There's a template available [on GitHub](https://github.com/LunarWatcher/se-data-dump-transformer/blob/master/config.example.json).
57
+
58
+ The only two fields you _need_ to fill out in the template is the email and password fields with credentials for a Stack Exchange account. You need to be logged in to download the data dumps, so the downloader needs the credentials to log in on your behalf. It doesn't matter if you're logged into SE elsewhere, as Selenium automatically creates a blank profile every time it starts, which won't include any cookies from SE, which means login is required.
59
+
60
+ > [!tip]
61
+ >
62
+ > The downloader can automatically create new accounts in the network for you, if you don't have all 180-whatever accounts on every site in the network already. You can also create these by hand if you prefer for some reason, but you are not required to have all 180+ accounts before using the downloader.
63
+
64
+ ## System requirements and pitfalls
65
+
66
+ `sedd` is exclusively Firefox-based, due to Chromium completely gutting support for uBlock Origin and custom filters. You need Firefox installed on your system to use `sedd`.
67
+
68
+ > [!note]
69
+ > On Linux and Windows-based systems, geckodriver is [slightly modified](https://pypi.org/project/undetected-geckodriver-lw/). This is an anti-anti-bot measure meant to prevent Cloudflare loops. If you're on macOS and get sent in a captcha loop, it's recommended you switch to Windows or Linux - a Linux VM is also an option if you have no way out of Apple's closed-down ecosystem.
70
+
71
+ Note that Ubuntu users, or other people who (for whatever reason) choose to use the Snap version of Firefox, have to jump through some extra hoops. The native version of Firefox is strongly encouraged, but if you run into problems with the snap version of Firefox and can't or won't switch, you need to define `export SE_GECKODRIVER=/snap/bin/geckodriver`. Selenium can and will find the snap version of `geckodriver` on its own, but for reasons I simply don't understand, it will still fail with several arbitrary errors.
72
+
73
+ ### Cloudflare issues or download issues.
74
+
75
+ Stack Exchange has configured Cloudflare to be _highly_ aggressive, especially to certain countries. You will almost certainly run into captchas, and the downloader is designed to deal with this. After an initial attempt to solve the captcha on its own, you'll be notified (provided you don't disable the notification provider in `config.json`) and asked to solve it manually.
76
+
77
+ If, at this point, it appears to succeed, but you're redirected back to a full-screen Cloudflare captcha wall, you've likely run into a Cloudflare loop. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cloudflare-loops) for further help. If this doesn't help, please open an issue.
78
+
79
+ If the downloads start fine, but later suddenly fail for no good reason, you're likely running into general download instability. This especially applies to `stackoverflow.com.7z`, as its massive size simply increases the chance you wait for it long enough that it flakes out. See [the main README](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#download-instability-particularly-of-stackoverflowcom7z) for further help.
80
+
81
+ The "Warnings" section in the README may contain additional information about other failure modes not listed here in the future.
82
+
83
+ ## Using the downloader
84
+
85
+ With `./config.json` in the current working directory and Firefox installed, you can now run the downloader with:
86
+ ```python3
87
+ sedd
88
+ ```
89
+
90
+ For command line flags, see `sedd --help`, or [the main readme](https://github.com/LunarWatcher/se-data-dump-transformer/tree/master?tab=readme-ov-file#cli-options).
91
+
@@ -0,0 +1,15 @@
1
+ LICENSE
2
+ README-python.md
3
+ README.md
4
+ pyproject.toml
5
+ sedd/__init__.py
6
+ sedd/__main__.py
7
+ sedd/cli.py
8
+ sedd/driver.py
9
+ sedd/main.py
10
+ sedd/utils.py
11
+ sedd.egg-info/PKG-INFO
12
+ sedd.egg-info/SOURCES.txt
13
+ sedd.egg-info/dependency_links.txt
14
+ sedd.egg-info/requires.txt
15
+ sedd.egg-info/top_level.txt
@@ -0,0 +1,5 @@
1
+ selenium==4.32.0
2
+ undetected-geckodriver-lw>=2.1.0
3
+ desktop-notifier==5.0.1
4
+ watchdog==4.0.2
5
+ requests
@@ -0,0 +1 @@
1
+ sedd
sedd-0.0.0/setup.cfg ADDED
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+