rssfeed 0.4.0__tar.gz → 0.4.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
rssfeed-0.4.2/PKG-INFO ADDED
@@ -0,0 +1,120 @@
1
+ Metadata-Version: 2.1
2
+ Name: rssfeed
3
+ Version: 0.4.2
4
+ Summary: A simple rss/atom feed parser
5
+ Keywords: rssfeed,rss,feed,atom,opml
6
+ Author: p7e4
7
+ License: GPLv3
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: Programming Language :: Python :: 3.12
10
+ Classifier: Programming Language :: Python :: 3.13
11
+ Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
12
+ Classifier: Topic :: Text Processing :: Markup :: XML
13
+ Project-URL: Repository, https://github.com/p7e4/rssfeed
14
+ Project-URL: Issues, https://github.com/p7e4/rssfeed/issues
15
+ Requires-Python: >=3.12
16
+ Requires-Dist: python-dateutil>=2.9
17
+ Description-Content-Type: text/markdown
18
+
19
+ # rssfeed
20
+
21
+ A simple rss/atom/opml parser
22
+
23
+ ## Installation
24
+
25
+ `pip install rssfeed`
26
+
27
+
28
+ ## Get Started
29
+
30
+ ### rss parse
31
+
32
+ ``` python
33
+ import requests
34
+ import rssfeed
35
+
36
+ text = requests.get("https://lobste.rs/rss").text
37
+ rssfeed.parse(text)
38
+ ```
39
+
40
+ ``` json
41
+ {
42
+ "name": "Lobsters",
43
+ "lastupdate": 1739824193,
44
+ "items": [
45
+ {
46
+ "title": "Why I'm Writing a Scheme Implementation in 2025 (The Answer is Async Rust)",
47
+ "author": "maplant.com by mplant",
48
+ "timestamp": 1739824193,
49
+ "url": "https://maplant.com/2025-02-17-Why-I'm-Writing-a-Scheme-Implementation-in-2025-(The-Answer-is-Async-Rust).html",
50
+ "content": "<p><a href=\"https://lobste.rs/s/zm1g8r/why_i_m_writing_scheme_implementation\">Comments</a></p>"
51
+ },
52
+ {
53
+ "title": "14 years of systemd",
54
+ "author": "lwn.net via calvin",
55
+ "timestamp": 1739814564,
56
+ "url": "https://lwn.net/SubscriberLink/1008721/7c31808d76480012/",
57
+ "content": "<p><a href=\"https://lobste.rs/s/c6rk0l/14_years_systemd\">Comments</a></p>"
58
+ },
59
+ {
60
+ "title": "Making the Web More Readable With Stylus",
61
+ "author": "wezm.net by wezm",
62
+ "timestamp": 1739757928,
63
+ "url": "https://www.wezm.net/v2/posts/2025/stylus/",
64
+ "content": "<p><a href=\"https://lobste.rs/s/sag0p3/making_web_more_readable_with_stylus\">Comments</a></p>"
65
+ }
66
+ ]
67
+ }
68
+ ```
69
+
70
+ > rssfeed **does not** escape HTML tags, which means you had to sanitization content otherwise it may lead to [Cross-site scripting](https://developer.mozilla.org/en-US/docs/Glossary/Cross-site_scripting) attacks, a recommended choice is [nh3](https://github.com/messense/nh3).
71
+
72
+
73
+ ### opml parse
74
+
75
+ ``` python
76
+ import rssfeed
77
+ opml = """
78
+ <?xml version="1.0" encoding="UTF-8"?>
79
+ <opml version="1.0">
80
+ <head>
81
+ <title>demo feeds</title>
82
+ </head>
83
+ <body>
84
+ <outline text="news">
85
+ <outline text="奇客Solidot" xmlUrl="https://www.solidot.org/index.rss" type="rss" />
86
+ <outline text="news-submenu">
87
+ <outline text="Lobsters" xmlUrl="https://lobste.rs/rss" type="rss" />
88
+ </outline>
89
+ </outline>
90
+ <outline text="阮一峰的网络日志" xmlUrl="https://feeds.feedburner.com/ruanyifeng" type="rss" />
91
+ </body>
92
+ </opml>
93
+ """
94
+ rssfeed.opmlParse(opml)
95
+ ````
96
+
97
+ ``` json
98
+ {
99
+ "default": [
100
+ {
101
+ "name": "阮一峰的网络日志",
102
+ "url": "https://feeds.feedburner.com/ruanyifeng"
103
+ }
104
+ ],
105
+ "news": [
106
+ {
107
+ "name": "奇客Solidot",
108
+ "url": "https://www.solidot.org/index.rss"
109
+ },
110
+ {
111
+ "name": "Lobsters",
112
+ "url": "https://lobste.rs/rss"
113
+ }
114
+ ]
115
+ }
116
+ ```
117
+
118
+
119
+
120
+
@@ -0,0 +1,102 @@
1
+ # rssfeed
2
+
3
+ A simple rss/atom/opml parser
4
+
5
+ ## Installation
6
+
7
+ `pip install rssfeed`
8
+
9
+
10
+ ## Get Started
11
+
12
+ ### rss parse
13
+
14
+ ``` python
15
+ import requests
16
+ import rssfeed
17
+
18
+ text = requests.get("https://lobste.rs/rss").text
19
+ rssfeed.parse(text)
20
+ ```
21
+
22
+ ``` json
23
+ {
24
+ "name": "Lobsters",
25
+ "lastupdate": 1739824193,
26
+ "items": [
27
+ {
28
+ "title": "Why I'm Writing a Scheme Implementation in 2025 (The Answer is Async Rust)",
29
+ "author": "maplant.com by mplant",
30
+ "timestamp": 1739824193,
31
+ "url": "https://maplant.com/2025-02-17-Why-I'm-Writing-a-Scheme-Implementation-in-2025-(The-Answer-is-Async-Rust).html",
32
+ "content": "<p><a href=\"https://lobste.rs/s/zm1g8r/why_i_m_writing_scheme_implementation\">Comments</a></p>"
33
+ },
34
+ {
35
+ "title": "14 years of systemd",
36
+ "author": "lwn.net via calvin",
37
+ "timestamp": 1739814564,
38
+ "url": "https://lwn.net/SubscriberLink/1008721/7c31808d76480012/",
39
+ "content": "<p><a href=\"https://lobste.rs/s/c6rk0l/14_years_systemd\">Comments</a></p>"
40
+ },
41
+ {
42
+ "title": "Making the Web More Readable With Stylus",
43
+ "author": "wezm.net by wezm",
44
+ "timestamp": 1739757928,
45
+ "url": "https://www.wezm.net/v2/posts/2025/stylus/",
46
+ "content": "<p><a href=\"https://lobste.rs/s/sag0p3/making_web_more_readable_with_stylus\">Comments</a></p>"
47
+ }
48
+ ]
49
+ }
50
+ ```
51
+
52
+ > rssfeed **does not** escape HTML tags, which means you had to sanitization content otherwise it may lead to [Cross-site scripting](https://developer.mozilla.org/en-US/docs/Glossary/Cross-site_scripting) attacks, a recommended choice is [nh3](https://github.com/messense/nh3).
53
+
54
+
55
+ ### opml parse
56
+
57
+ ``` python
58
+ import rssfeed
59
+ opml = """
60
+ <?xml version="1.0" encoding="UTF-8"?>
61
+ <opml version="1.0">
62
+ <head>
63
+ <title>demo feeds</title>
64
+ </head>
65
+ <body>
66
+ <outline text="news">
67
+ <outline text="奇客Solidot" xmlUrl="https://www.solidot.org/index.rss" type="rss" />
68
+ <outline text="news-submenu">
69
+ <outline text="Lobsters" xmlUrl="https://lobste.rs/rss" type="rss" />
70
+ </outline>
71
+ </outline>
72
+ <outline text="阮一峰的网络日志" xmlUrl="https://feeds.feedburner.com/ruanyifeng" type="rss" />
73
+ </body>
74
+ </opml>
75
+ """
76
+ rssfeed.opmlParse(opml)
77
+ ````
78
+
79
+ ``` json
80
+ {
81
+ "default": [
82
+ {
83
+ "name": "阮一峰的网络日志",
84
+ "url": "https://feeds.feedburner.com/ruanyifeng"
85
+ }
86
+ ],
87
+ "news": [
88
+ {
89
+ "name": "奇客Solidot",
90
+ "url": "https://www.solidot.org/index.rss"
91
+ },
92
+ {
93
+ "name": "Lobsters",
94
+ "url": "https://lobste.rs/rss"
95
+ }
96
+ ]
97
+ }
98
+ ```
99
+
100
+
101
+
102
+
@@ -19,15 +19,16 @@ keywords = [
19
19
  "rss",
20
20
  "feed",
21
21
  "atom",
22
- "feedparser",
22
+ "opml",
23
23
  ]
24
24
  classifiers = [
25
25
  "Programming Language :: Python :: 3",
26
26
  "Programming Language :: Python :: 3.12",
27
+ "Programming Language :: Python :: 3.13",
27
28
  "License :: OSI Approved :: GNU General Public License v3 (GPLv3)",
28
29
  "Topic :: Text Processing :: Markup :: XML",
29
30
  ]
30
- version = "0.4.0"
31
+ version = "0.4.2"
31
32
 
32
33
  [project.license]
33
34
  text = "GPLv3"
@@ -0,0 +1,3 @@
1
+ from .lib import parse, opmlParse, ParseError
2
+
3
+ __all__ = ["parse", "opmlParse", "ParseError"]
@@ -1,22 +1,26 @@
1
1
  from dateutil.parser import parse as timeParse
2
2
  from xml.etree import ElementTree
3
3
 
4
- __version__ = "0.4.0"
4
+ __version__ = "0.4.2"
5
5
 
6
6
  class ParseError(Exception):
7
7
  pass
8
8
 
9
- def parse(data):
9
+ def _parse(data):
10
+ if not (data:=data.lstrip()):
11
+ raise ParseError("empty data")
10
12
  parser = ElementTree.XMLPullParser(("start", "end"))
11
13
  try:
12
14
  parser.feed(data)
13
15
  parser.close()
14
16
  except ElementTree.ParseError as e:
15
17
  raise ParseError("xml parse fail") from e
18
+ return parser
16
19
 
20
+ def parse(data, url=None):
21
+ if url: url = url[:8] + url[8:].rsplit("/")[0]
17
22
  items = list()
18
- lastTag = str()
19
- for event, elem in parser.read_events():
23
+ for event, elem in _parse(data).read_events():
20
24
  tag = elem.tag.split("}", 1)[1] if elem.tag.startswith("{") else elem.tag
21
25
  text = elem.text.strip() if elem.text else str()
22
26
  if event == "start":
@@ -31,9 +35,9 @@ def parse(data):
31
35
  else:
32
36
  i = items[-1]
33
37
  match tag:
34
- case "content" | "description" | "encoded" | "summary":
38
+ case "description" | "encoded" | "summary" | "content":
35
39
  i["content"] = text
36
- case "updated" | "pubDate" | "published" | "lastBuildDate":
40
+ case "pubDate" | "updated" | "published" | "lastBuildDate":
37
41
  if text.isdigit():
38
42
  i["timestamp"] = int(text)
39
43
  elif text:
@@ -43,13 +47,11 @@ def parse(data):
43
47
  raise ParseError("time parse fail") from e
44
48
  case "link":
45
49
  i["url"] = text or elem.get("href")
46
- case "name" if lastTag == "author":
47
- i["author"] = text
48
- case "title":
50
+ if url and not i["url"].startswith("http"):
51
+ i["url"] = f"{url}/{i["url"].lstrip("/")}"
52
+ case "title" | "author":
49
53
  i[tag] = text
50
54
 
51
- lastTag = tag
52
-
53
55
  if not items:
54
56
  raise ParseError("not valid result")
55
57
 
@@ -60,3 +62,28 @@ def parse(data):
60
62
  }
61
63
 
62
64
  return feed
65
+
66
+ def opmlParse(data):
67
+ path = list()
68
+ result = dict(default=list())
69
+ for event, elem in _parse(data).read_events():
70
+ if elem.tag != "outline":
71
+ continue
72
+ if event == "start":
73
+ if elem.get("type") == "rss":
74
+ name = path[0] if path else "default"
75
+ result[name].append({
76
+ "name": elem.get("text") or elem.get("title"),
77
+ "url": elem.get("xmlUrl") or elem.get("htmlUrl")
78
+ })
79
+ else:
80
+ path.append(elem.get("text") or elem.get("title"))
81
+ if len(path) == 1: result[path[0]] = list()
82
+ else:
83
+ if elem.get("type") != "rss":
84
+ path.pop()
85
+
86
+ if not result["default"]:
87
+ del result["default"]
88
+
89
+ return result
rssfeed-0.4.0/PKG-INFO DELETED
@@ -1,73 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: rssfeed
3
- Version: 0.4.0
4
- Summary: A simple rss/atom feed parser
5
- Keywords: rssfeed,rss,feed,atom,feedparser
6
- Author: p7e4
7
- License: GPLv3
8
- Classifier: Programming Language :: Python :: 3
9
- Classifier: Programming Language :: Python :: 3.12
10
- Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
11
- Classifier: Topic :: Text Processing :: Markup :: XML
12
- Project-URL: Repository, https://github.com/p7e4/rssfeed
13
- Project-URL: Issues, https://github.com/p7e4/rssfeed/issues
14
- Requires-Python: >=3.12
15
- Requires-Dist: python-dateutil>=2.9
16
- Description-Content-Type: text/markdown
17
-
18
- # rssfeed
19
-
20
- A simple rss/atom feed parser
21
-
22
- ## Installation
23
-
24
- `pip install rssfeed`
25
-
26
- ## Get Started
27
-
28
- ``` python
29
- import requests
30
- import rssfeed
31
-
32
- feed = rssfeed.parse(requests.get("https://www.solidot.org/index.rss").text)
33
- print(feed)
34
- ```
35
- ```
36
- {
37
- "name": "奇客Solidot–传递最新科技情报",
38
- "lastupdate": 1717423475,
39
- "items": [
40
- {
41
- "title": "中国科学家使用细胞疗法治愈一名患者的糖尿病",
42
- "author": "",
43
- "timestamp": 1717410594,
44
- "url": "https://www.solidot.org/story?sid=78338",
45
- "content": "《南华早报》报道,中国科学家利用细胞疗法成功治愈了一名患者的糖尿病。研究报告发表在《Cell Discovery》期刊 ..."
46
- },
47
- {
48
- "title": "Steam 平台 Linux 玩家四分之三使用 AMD CPU",
49
- "author": "",
50
- "timestamp": 1717404736,
51
- "url": "https://www.solidot.org/story?sid=78337",
52
- "content": "根据 Valve 公布的 Steam 硬件和软件调查,Linux 份额在过去的五月增长了 0.42% 至 2.32%,macOS 增至 1.47% ..."
53
- },
54
- {
55
- "title": "Hugging Face 称黑客窃取了 Spaces 平台的身份验证令牌",
56
- "author": "",
57
- "timestamp": 1717400574,
58
- "url": "https://www.solidot.org/story?sid=78336",
59
- "content": "Hugging Face 官方博客披露黑客窃取了其 Spaces 平台的身份验证令牌。Spaces 是社区用户创建和递交 AI 应用的库 ..."
60
- }
61
- ...
62
- ]
63
- }
64
- ```
65
-
66
- ## Warning
67
-
68
- rssfeed **does not** escape any HTML tags, which mean if you does not check the content and display it somewhere html can be rendered, it may lead to [Cross-site scripting](https://developer.mozilla.org/en-US/docs/Glossary/Cross-site_scripting) attacks.
69
-
70
- ## Changelog
71
-
72
- [Changelog.md](/Changelog.md)
73
-
rssfeed-0.4.0/README.md DELETED
@@ -1,56 +0,0 @@
1
- # rssfeed
2
-
3
- A simple rss/atom feed parser
4
-
5
- ## Installation
6
-
7
- `pip install rssfeed`
8
-
9
- ## Get Started
10
-
11
- ``` python
12
- import requests
13
- import rssfeed
14
-
15
- feed = rssfeed.parse(requests.get("https://www.solidot.org/index.rss").text)
16
- print(feed)
17
- ```
18
- ```
19
- {
20
- "name": "奇客Solidot–传递最新科技情报",
21
- "lastupdate": 1717423475,
22
- "items": [
23
- {
24
- "title": "中国科学家使用细胞疗法治愈一名患者的糖尿病",
25
- "author": "",
26
- "timestamp": 1717410594,
27
- "url": "https://www.solidot.org/story?sid=78338",
28
- "content": "《南华早报》报道,中国科学家利用细胞疗法成功治愈了一名患者的糖尿病。研究报告发表在《Cell Discovery》期刊 ..."
29
- },
30
- {
31
- "title": "Steam 平台 Linux 玩家四分之三使用 AMD CPU",
32
- "author": "",
33
- "timestamp": 1717404736,
34
- "url": "https://www.solidot.org/story?sid=78337",
35
- "content": "根据 Valve 公布的 Steam 硬件和软件调查,Linux 份额在过去的五月增长了 0.42% 至 2.32%,macOS 增至 1.47% ..."
36
- },
37
- {
38
- "title": "Hugging Face 称黑客窃取了 Spaces 平台的身份验证令牌",
39
- "author": "",
40
- "timestamp": 1717400574,
41
- "url": "https://www.solidot.org/story?sid=78336",
42
- "content": "Hugging Face 官方博客披露黑客窃取了其 Spaces 平台的身份验证令牌。Spaces 是社区用户创建和递交 AI 应用的库 ..."
43
- }
44
- ...
45
- ]
46
- }
47
- ```
48
-
49
- ## Warning
50
-
51
- rssfeed **does not** escape any HTML tags, which mean if you does not check the content and display it somewhere html can be rendered, it may lead to [Cross-site scripting](https://developer.mozilla.org/en-US/docs/Glossary/Cross-site_scripting) attacks.
52
-
53
- ## Changelog
54
-
55
- [Changelog.md](/Changelog.md)
56
-
@@ -1 +0,0 @@
1
- from .lib import parse
File without changes