mirror of
https://git.notmuchmail.org/git/notmuch
synced 2024-11-24 20:08:10 +01:00
test: add known broken test for indexing html
'quite' on IRC reported that notmuch new was grinding to a halt during initial indexing, and we eventually narrowed the problem down to some html parts with large embedded images. These cause the number of terms added to the Xapian database to explode (the first 400 messages generated 4.6M unique terms), and of course the resulting terms are not much use for searching. The second test is sanity check for any "improved" indexing of HTML.
This commit is contained in:
parent
e565118172
commit
77c9ec1fdd
4 changed files with 106 additions and 0 deletions
19
test/T680-html-indexing.sh
Executable file
19
test/T680-html-indexing.sh
Executable file
|
@ -0,0 +1,19 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
test_description="indexing of html parts"
|
||||||
|
. ./test-lib.sh || exit 1
|
||||||
|
|
||||||
|
add_email_corpus html
|
||||||
|
|
||||||
|
test_begin_subtest 'embedded images should not be indexed'
|
||||||
|
test_subtest_known_broken
|
||||||
|
notmuch search kwpza7svrgjzqwi8fhb2msggwtxtwgqcxp4wbqr4wjddstqmeqa7 > OUTPUT
|
||||||
|
test_expect_equal_file /dev/null OUTPUT
|
||||||
|
|
||||||
|
test_begin_subtest 'non tag text should be indexed'
|
||||||
|
notmuch search hunter2 | notmuch_search_sanitize > OUTPUT
|
||||||
|
cat <<EOF > EXPECTED
|
||||||
|
thread:XXX 2009-11-17 [1/1] David Bremner; test html attachment (inbox unread)
|
||||||
|
EOF
|
||||||
|
test_expect_equal_file EXPECTED OUTPUT
|
||||||
|
|
||||||
|
test_done
|
|
@ -9,3 +9,6 @@ default
|
||||||
broken
|
broken
|
||||||
The broken corpus contains messages that are broken and/or RFC
|
The broken corpus contains messages that are broken and/or RFC
|
||||||
non-compliant, ensuring we deal with them in a sane way.
|
non-compliant, ensuring we deal with them in a sane way.
|
||||||
|
|
||||||
|
html
|
||||||
|
The html corpus contains html parts
|
||||||
|
|
15
test/corpora/html/attribute-text
Normal file
15
test/corpora/html/attribute-text
Normal file
|
@ -0,0 +1,15 @@
|
||||||
|
From: David Bremner <david@example.net>
|
||||||
|
To: David Bremner <david@example.net>
|
||||||
|
Subject: test html attachment
|
||||||
|
Date: Tue, 17 Nov 2009 21:28:38 +0600
|
||||||
|
Message-ID: <87d1dajhgf.fsf@example.net>
|
||||||
|
MIME-Version: 1.0
|
||||||
|
Content-Type: text/html
|
||||||
|
Content-Disposition: inline; filename=test.html
|
||||||
|
|
||||||
|
<html>
|
||||||
|
<body>
|
||||||
|
<input value="a>swordfish">
|
||||||
|
</body>
|
||||||
|
hunter2
|
||||||
|
</html>
|
69
test/corpora/html/embedded-image
Normal file
69
test/corpora/html/embedded-image
Normal file
|
@ -0,0 +1,69 @@
|
||||||
|
From: =?utf-8?b?bWFsbW9ib3Jn?= <daemon@lublin.se>
|
||||||
|
To: =?utf-8?b?Ym9lbmRlLm1hbG1vYm9yZw==?= <daemon@lublin.se>
|
||||||
|
Date: Tue, 19 Jul 2016 11:54:24 +0200
|
||||||
|
X-Feed2Imap-Version: 1.2.5
|
||||||
|
Message-Id: <boendemalmoborg-1834@eltanin.uberspace.de>
|
||||||
|
Subject: =?utf-8?b?VGFjayBhbGxhIHRyYWZpa2FudGVyIG9jaCBmb3Rnw6RuZ2FyZSE=?=
|
||||||
|
Content-Type: multipart/alternative; boundary="=-1468922508-176605-12427-9500-21-="
|
||||||
|
MIME-Version: 1.0
|
||||||
|
|
||||||
|
|
||||||
|
--=-1468922508-176605-12427-9500-21-=
|
||||||
|
Content-Type: text/plain; charset=utf-8; format=flowed
|
||||||
|
Content-Transfer-Encoding: 8bit
|
||||||
|
|
||||||
|
<http://malmoborg.se/2016/07/tack-alla-trafikanter-och-fotgangare/>
|
||||||
|
|
||||||
|
Malmö 2016-07-09
|
||||||
|
|
||||||
|
I skrivande stund är vi i färd med att avetablera vår entreprenad på
|
||||||
|
Tigern 3, Regementsgatan 6 i Malmö. Fastigheten har genomgått ett större
|
||||||
|
dräneringsarbete som i sin tur har inneburit vissa
|
||||||
|
trafikbegränsningar på Regementsgatan samt Davidshallsgatan under några
|
||||||
|
veckors tid. Fastighetsägaren är mycket nöjd med vår arbetsinsats och vi
|
||||||
|
kan glatt meddela att båda vägfilerna kommer att öppnas inom kort. Nu
|
||||||
|
kommer den vackra fastigheten att klara sig torrskodd under många år
|
||||||
|
framöver [A]
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
[A] http://malmoborg.se/wp-includes/images/smilies/icon_smile.gif
|
||||||
|
--
|
||||||
|
Feed: Förvaltnings AB Malmöborg
|
||||||
|
<http://malmoborg.se>
|
||||||
|
Item: Tack alla trafikanter och fotgängare!
|
||||||
|
<http://malmoborg.se/2016/07/tack-alla-trafikanter-och-fotgangare/>
|
||||||
|
Date: 2016-07-19 11:54:24 +0200
|
||||||
|
Author: malmoborg
|
||||||
|
Filed under: Nyheter
|
||||||
|
|
||||||
|
--=-1468922508-176605-12427-9500-21-=
|
||||||
|
Content-Type: text/html; charset=utf-8
|
||||||
|
Content-Transfer-Encoding: 8bit
|
||||||
|
|
||||||
|
<table border="1" width="100%" cellpadding="0" cellspacing="0" borderspacing="0"><tr><td>
|
||||||
|
<table width="100%" bgcolor="#EDEDED" cellpadding="4" cellspacing="2">
|
||||||
|
<tr><td align="right"><b>Feed:</b></td>
|
||||||
|
<td width="100%"><a href="http://malmoborg.se">
|
||||||
|
<b>Förvaltnings AB Malmöborg</b>
|
||||||
|
</a>
|
||||||
|
</td></tr><tr><td align="right"><b>Item:</b></td>
|
||||||
|
<td width="100%"><a href="http://malmoborg.se/2016/07/tack-alla-trafikanter-och-fotgangare/"><b>Tack alla trafikanter och fotgängare!</b>
|
||||||
|
</a>
|
||||||
|
</td></tr></table></td></tr></table>
|
||||||
|
|
||||||
|
<p>Malmö 2016-07-09</p>
|
||||||
|
<p>I skrivande stund är vi i färd med att avetablera vår entreprenad på Tigern 3, Regementsgatan 6 i Malmö. Fastigheten har genomgått ett större dräneringsarbete som i sin tur har inneburit vissa trafikbegränsningar på Regementsgatan samt Davidshallsgatan under några veckors tid. Fastighetsägaren är mycket nöjd med vår arbetsinsats och vi kan glatt meddela att båda vägfilerna kommer att öppnas inom kort. Nu kommer den vackra fastigheten att klara sig torrskodd under många år framöver <img src="data:image/gif;base64,R0lGODlhDwAPALMOAP/qAEVFRQAAAP/OAP/JAP+0AP6dAP/+k//9E///////
|
||||||
|
xzMzM///6//lAAAAAAAAACH5BAEAAA4ALAAAAAAPAA8AAARb0EkZap3YVabO
|
||||||
|
GRcWcAgCnIMRTEEnCCfwpqt2mHEOagoOnz+CKnADxoKFyiHHBBCSAdOiCVg8
|
||||||
|
KwPZa7sVrgJZQWI8FhB2msGgwTXTWGqCXP4WBQr4wjDDstQmEQA7
|
||||||
|
" alt=":-)" class="wp-smiley" /> </p>
|
||||||
|
<p> </p>
|
||||||
|
<hr width="100%"/>
|
||||||
|
<table width="100%" cellpadding="0" cellspacing="0">
|
||||||
|
<tr><td align="right"><font color="#ababab">Date:</font> </td><td><font color="#ababab">2016-07-19 11:54:24 +0200</font></td></tr>
|
||||||
|
<tr><td align="right"><font color="#ababab">Author:</font> </td><td><font color="#ababab">malmoborg</font></td></tr>
|
||||||
|
<tr><td align="right"><font color="#ababab">Filed under:</font> </td><td><font color="#ababab">Nyheter</font></td></tr>
|
||||||
|
</table>
|
||||||
|
|
||||||
|
--=-1468922508-176605-12427-9500-21-=--
|
Loading…
Reference in a new issue