All posts in the topic Government PDFs, OCR etc. (Short link)
Summary
- There are 62 posts — by 18 authors — in this topic.
- Latest post made by gordon.b.anderson at Nov 22 02:14 NZDT
| From | File | Date |
|---|---|---|
| Jonathan Hunt | Hurunui_Tribunal_report_200908.pdf | Aug 16 15:44 NZST |
Hi all I just received the Special Tribunal report re a Water Conservation Order on the Hurunui as a PDF comprised of scanned pages. I've immediately replied to MfE asking if they will have an accessible text version available, and if so, when. I'm sure this isn't the first time. Such documents are problematic: - can't be easily scanned or analysed by searching on keywords. - can't be quoted via copy & paste - can't be parsed by search engines (incl. Google, local search systems (OSX Spotlight et al), web search). - can't be electronically cited or referenced. Some thoughts: - how prevalent is this practice? (Presumably, someone runs a printed doc through a page feeder on a photocopier, then emails out the PDF). - would it be acceptable if they OCR'ed and proofed the OCR output? - if Govt is unwilling/unable/ignorant to make accessible text, is there scope for crowd-sourcing this? I'm sure we could develop a system to: a. ingest scanned PDFs, b. break them into pages, c. queue pages for OCR, d. OCR each page (via web service, or open source OCR package), e. queue results for proofing to one or more volunteers, f. recompile document into scannable PDF (+ HTML, ODF for good measure). Is there any service like this already? Once the service was in place it would lend itself for being a general repository for government documents: a. unique, persistent document identifiers for all docs. b. citeability (e.g. http://citability.org/) c. comments (http://www.commentonthis.com/about/ ). d. search e. revision tracking... Hmm, needs some funding I think... Would be nicer to fix at source...
I'd love to know what became of the accessibility requirements I remember from
not too long ago, that required everything on .govt.nz websites to be available
in html so (amongst other things) they were readable by the blind using
screen-readers, and available for people on slow internet connections.
This seems to have slowly vanished into the ether... Does anyone know if it's
still alive?
You've just described Distributed Proofreaders - http://www.pgdp.net/ It increased the output of Project Gutenburg massively when it launched. Also reCaptcha for fixing hard to OCR pieces of text - http://recaptcha.net/
Isabella
I suspect a lot of that sort of stuff (including multi-browser support)
is disappearing as a result of the rush to Sharepoint by government
agencies. Also not helped by various movements and redundancies at SSC.
Cheers
Don
On Fri, 2009-08-14 at 16:27 +1200, <email obscured>
wrote:
> I'd love to know what became of the accessibility requirements I remember
from not too long ago, that required everything on .govt.nz websites to be
available in html so (amongst other things) they were readable by the blind
using screen-readers, and available for people on slow internet connections.
>
> This seems to have slowly vanished into the ether... Does anyone know if it's
still alive?
-timClicks 2009/8/14 <email obscured>> > Hi all > > I just received the Special Tribunal report re a Water Conservation Order > on the Hurunui as a PDF comprised of scanned pages. > > I've immediately replied to MfE asking if they will have an accessible > text version available, and if so, when. > > At source, this is most likely been given to an administrator who's just used their Konica to create a PDF for you. Unless they have Adobe Professional, they're likely to not have the ability to OCR it in-house. If they do have access to Adobe Professional, they might not know this functionality exists. If they do, they may think that it's a waste of time - since you'll have the information either way. Do you run Ubuntu at all? "sudo apt-get install pdfedit" may help: http://www.addictivetips.com/ubuntu-linux-tips/edit-pdf-document-in-ubuntu-linux-with-pdfedit/ > Some thoughts: > - how prevalent is this practice? (Presumably, someone runs a printed doc > through a page feeder on a photocopier, then emails out the PDF). > - would it be acceptable if they OCR'ed and proofed the OCR output? > - if Govt is unwilling/unable/ignorant to make accessible text, is there > scope for crowd-sourcing this? > > At my old job at MED, we pushed through *lots* of OIA requests. No one really seemed to care except me that people might want to be able to copy & paste text. That extra 2mins is valuable.
I'd love to know what became of the accessibility requirements I remember from
not too long ago, that required everything on .govt.nz websites to be available
in html so (amongst other things) they were readable by the blind using
screen-readers, and available for people on slow internet connections.
This seems to have slowly vanished into the ether... Does anyone know if it's
still alive?
On Fri, Aug 14, 2009 at 4:27 PM, <email obscured>> wrote: > I'd love to know what became of the accessibility requirements I remember > from not too long ago > They are still there: http://www.webstandards.govt.nz/applications-and-accessible-alternatives/
I'm not sure about the specific requirements of newzealand.govt.nz. Still, undoubtedly they'll want a high level of compliance with the overall govt web requirements http://www.webstandards.govt.nz/technical/. It requires NZ govt sites to meet "AA" compliance with the w3 Web Content Accessiblity Guidelines, http://www.w3.org/TR/WCAG20/. Basically, yes - things need to be text readable. The NZ standards were updated in March 09, so should still be fresh in CTOs' minds. As Don mentioned, departmental buy-in is agency specific. Without a core agency caning, some CEs may choose to do their own thing. 2009/8/14 Don <email obscured>>
> I just received the Special Tribunal report re a Water Conservation Order > on the Hurunui as a PDF comprised of scanned pages. > Sometimes PDF is used as a barrier to reuse. The UK Guardian just went through an exercise of crowdsourcing the interpretation of 100,000s PDF pages of politicians' expenses. http://mps-expenses.guardian.co.uk/. I note our government did the same (but don't blame conspiracy, when apathy will do). If you know someone with OmniPage Suite OCR, you can batch OCR a complete PDF. They also do as a SDK. http://www.nuance.com/imaging/products/omnipage.asp
> At my old job at MED, we pushed through *lots* of OIA requests. No one
> really seemed to care except me that people might want to be able to copy &
> paste text. That extra 2mins is valuable.
>
> Perhaps the answer is to make the Official Information Act more
user-centric.
Rather than users be charged for the cost of extracting the information,
departments are charged the cost of the users converting the provided
information to an accessible form. A cost incentive would encourage
greater consideration of efficiency/effectiveness in the beginning.
Wasn't there initially two things;
a) the pages were scanned implying that there was only a paper version
available (it was a Water Conservation Order which suggests that it is
verging on being an historic document created before the 1990s i.e. last
century ~8-) It was possibly type in triplicate! )
b) the images were sent as PDFs
Some things to note here are:
1. Paper is one of the most enduring media for information
2. There is a lot of information solely in paper form
3. The A4 format for information is very effective and efficient.
4. Government will be using paper for decades yet because most citizens
are comfortable and familiar with paper
5. The management of paper is well understood and supported by
bureaucracy
6. With paper you can quickly see the document size, "density" of
information, rate of progress, book marks, annotations etc
7. Paper requires very little power and connectivity.
8. Paper is far more convenient at meetings (no need for multiple power
plugs, projectors etc)
9. PDFs get rid of fingerprints or stray comments.
10. PDFs are easy to "clean"
11. Information and authority and how it is "presented" can be "secured"
in PDFs
12. PDFs "look and feel" like paper
13. PDF's can be treated much like paper by the bureaucracy
14. PDF's can be transported quicker than paper
Paper is far more end-user centric to most people.
Personally I find paper a pain, but it has some very desirable features.
As for PDFs - they bridge the paper / online world better than some other
formats.
Fixing at source has some superficial appeal, until we start thinking about
what it is we are trying to fix. Scanning and OCRing all paper is not
practical. As for new documents, the bureaucracy is making a start but there
are still many of the above "features" that need be addressed.
When the stuff they are sending you is printed copies of emails, you do
start to get suspicious about what's in the "track changes" settings...
> On 15/08/2009, at 5:53 PM, Jim McLeod wrote: > Wasn't there initially two things; > > a) the pages were scanned implying that there was only a paper version > available (it was a Water Conservation Order which suggests that it is > verging on being an historic document created before the 1990s i.e. last > century ~8-) It was possibly type in triplicate! ) It was a PDF of a scanned paper document. This document is dated 5 August 2009 (i.e. 10 days ago). The document is a scan of the formal recommendations from the Special Tribunal formed to consider a WCO on the Hurunui River. http://www.mfe.govt.nz/issues/water/freshwater/water-conservation/application-water-conservation.html > b) the images were sent as PDFs PDF is fine, but a PDF of a scanned image is poor practice unless steps are taken to make the information contained the document accessible. It's quite possible to make a PDF that appears (and prints) as a scanned paper document but contains the same text in machine-readable text suitable for copy and paste, search, etc. While paper has many fine attributes, it's expensive and inefficient to distribute. Also, paper doesn't lend itself to citeability, search, transclusion etc. A single canonical instance can be a common reference point for comments, annotations etc. becoming a communal information resource. IMHO, Government should be publishing canonical versions of all such documents online, identified in a common national (global) namespace, with full accessibility.
My mistake - not an old document. > The document is a scan of the formal recommendations from the Special > Tribunal formed to consider a WCO on the Hurunui River. <http://www.mfe.govt.nz/issues/water/freshwater/water-conservation/application-water-conservation.html> > http://www.mfe.govt.nz/issues/ > water/freshwater/water-conservation/application-water-conservation.html However, not sure which document you are referring to amongst the stack at the end of the link. Rather than a single document, I see that many documents have been gathered and that much information has to be considered. This situation is not unique as it occurs in many formal processes. The relatively recent step of making the documents available in electronic form via a Web site is significant progress. That apparently they are all in PDF format is at least consistent. And doesn't this provide the start of the canonical, online, repository you seek? In the sample PDFs I grabbed text, searched it and copied and pasted it else where. Many of the documents from expert witnesses likely cite other documents. However, I do agree that with technology there is the potential to do far better. For example, it would be very helpful to be able to access the actual data behind graphs, to have interactive maps of the geographic areas, to be able to identify themes both common and uncommon across all documents (including transcripts of verbal evidence and cross examination and photos/images) ... In other words to sort out the evidence so it is available to help replace opinions in the decision making. One option I have been exploring is to accept that though the PDFs are not ideal at least they are available and to attempt to work with them. My default toolset includes Google apps ("free", online, relatively easy to use, continually expanding e.g. Wave), however such tools do require good connectivity and that isn't always available. (Yes we could focus on changing the process or we can use technology to explore, extend and improve the context around the present process.) Having got and assimilated the data and information the next challenge is effectively and efficiently engaging with the decision making process. The big hurdle I think for most people is the emphasis placed upon the physical meeting. In some areas we are starting to see aspects of telemeeting, but those physically in the meeting room so far tend to hold the greatest influence on the decisions. Perhaps we should be doing more to explore and encourage virtual meetings, such as in 2nd life, over physical meetings. Again a big challenge is connectivity.
On Sun, August 16, 2009 2:34 pm, Jim McLeod wrote: > My mistake - not an old document. > >> The document is a scan of the formal recommendations from the Special >> Tribunal formed to consider a WCO on the Hurunui River. > > <http://www.mfe.govt.nz/issues/water/freshwater/water-conservation/application-water-conservation.html> >> http://www.mfe.govt.nz/issues/ >> water/freshwater/water-conservation/application-water-conservation.html. > However, not sure which document you are referring to amongst the stack at > the end of the link. > Rather than a single document, I see that many documents have been > gathered > and that much information has to be considered. Those pages are for background only. The Tribunal report is not yet available online (another issue for another day). I've attached the specific document I received from MfE so you can see the issue directly > This situation is not unique as it occurs in many formal processes. The > relatively recent step of making the documents available in electronic > form > via a Web site is significant progress. That apparently they are all in > PDF > format is at least consistent. And doesn't this provide the start of the > canonical, online, repository you seek? In the sample PDFs I grabbed text, > searched it and copied and pasted it else where. Many of the documents > from > expert witnesses likely cite other documents. It falls well short of a canonical document repository. These documents are local to MfE. There's no consistency of document naming, metadata, etc. Many of the PDFs of evidence submitted are still scanned images. > However, I do agree that with technology there is the potential to do far > better. For example, it would be very helpful to be able to access the > actual data behind graphs, to have interactive maps of the geographic > areas, > to be able to identify themes both common and uncommon across all > documents > (including transcripts of verbal evidence and cross examination and > photos/images) ... In other words to sort out the evidence so it is > available to help replace opinions in the decision making. I agree. I understand it will take time, but I'd like to know that Government was moving in this direction. Interactive maps are not required, as long as geographically relevant data is marked up properly. > One option I have been exploring is to accept that though the PDFs are not > ideal at least they are available and to attempt to work with them. Please note that my issue is not with PDFs as such; it is with PDFs of scanned documents where no attempt has apparently been made to OCR the scanned text and thus make the contents available via copy/paste, search etc.
The following file was added to this topic:
>
>
> Those pages are for background only. The Tribunal report is not yet
> available online (another issue for another day).
> I've attached the specific document I received from MfE so you can see the
> issue directly
What you have here is the image of the actual genuine paper document signed
by the Tribunal members. (You might even see their finger prints on the
document ~8- [?]) ) That's why it is a PDF of an image. There are many
opportunities to get the wrong version as multiple people edit these things
both electronically and on paper. So it is good to get the signed one - each
page initialled as well to show that it has been "read". You might even get
last minute changes made in pen on the paper and they would also get their
specific individual legal signature. This is the legal process and the legal
document.
>
> It falls well short of a canonical document repository. These documents
> are local to MfE. There's no consistency of document naming, metadata,
> etc. Many of the PDFs of evidence submitted are still scanned images.
I admit that I haven't pulled the legal description of a canonical document
repository, but this is likely to have been a hearing in a physical meeting
where relevant documents are on paper and are physically handed to the
appropriate person for consideration by the Tribunal. Some of the Tribunal
may have documents on computers but most likely not. The papers can be
annotated at any time prior to being handed to the official. And can also be
modified during the submission. Hence, the "record" will be of these scanned
paper documents.
> I agree. I understand it will take time, but I'd like to know that
> Government was moving in this direction. Interactive maps are not
> required, as long as geographically relevant data is marked up properly.
People in government are trying to move things but need to also be able to
give very conservative participants confidence that the legal integrity of
the process is being "protected".
Sorry I didn't mean to pre-empt the app for displaying the geospatial data,
and agree the actual data is the objective. However, at present geospatial
data as with graphs excreta are submitted in paper form. In some cases this
can be in A0 or larger formats across multiple sheets. This does create a
challenge when attempting to construct an electronic record.
>
> > One option I have been exploring is to accept that though the PDFs are
> not
> > ideal at least they are available and to attempt to work with them.
>
> Please note that my issue is not with PDFs as such; it is with PDFs of
> scanned documents where no attempt has apparently been made to OCR the
> scanned text and thus make the contents available via copy/paste, search
> etc.
The challenge here is you then get an interpretation of the legal document.
Lawyers make much hay whenever legal documents are "interpreted". Recall
these things a primarily paper based processes, with all the opportunity to
annotate and change the words right up to the point of sign-off.
>
> I agree that the decision making process still needs plenty of work. But
> we are not out of the data woods yet! Not while Government emits PDFs of
> scanned images, or proprietary data formats, or encumbers raw data in
> non-helpful web UIs, or doesn't even make the raw data available.
Correct, but I think we have to be aware that the actual Tribunal hearing is
still a paper based legal process and this will force the generation of
paper.
> My concern is that by not publishing accessible data, Government
> disenfranchises many citizens who don't even get to the decision-making
> stage (as a participant) because they never found out the decision was in
> play. It's amazing the number of agencies who feel that publishing a
> public notice in a regional newspaper is legitimate notification.
I am reasonably sure that a public notice in a local newspaper *is* the
legally legitimate notification and has been for many decades. However, I do
know many organisations attempt to do more than this. Their challenge is
where is an effective and efficient place to put a notification? Though more
people are starting to engage online many (a majority?) still don't.
> Instead, if citizens can rely on authoritative, timely online publication
> of notices, consultations, legislation etc. then citizens can start to
> build the tools to parse, analyse, track, notify, annotate etc. the stream
> of government activity most relevant to them.
Hence the importance of Government RSS feeds, metadata standards, open
data formats, microformats, etc.
I am converted, but we still have a large rump who haven't and possibly
never will. And because of them we will be supporting cumbersome paper
systems for quite some time yet.
My present challenge is to understand what happened in a string of "courts"
over the last couple of decades where tonnes of paper has been generated and
decisions made. Once I have a grip on this then I have to get my head around
what it means for the actual "management" of the natural resource.
On Sun, August 16, 2009 5:56 pm, Jim McLeod wrote:
> What you have here is the image of the actual genuine paper document
> signed
> by the Tribunal members. (You might even see their finger prints on the
> document ~8- [?]) ) That's why it is a PDF of an image. There are many
> opportunities to get the wrong version as multiple people edit these
> things
> both electronically and on paper. So it is good to get the signed one -
> each
> page initialled as well to show that it has been "read". You might even
> get
> last minute changes made in pen on the paper and they would also get their
> specific individual legal signature. This is the legal process and the
> legal
> document.
*sigh* I am aware that it is a legal document, and that a scanned image is
'potentially' the most likely to be 'valid'.
What I am complaining about is that a document comprising scanned images
is useless except for reading by a human, unless it also includes
accessible text (i.e. by OCR).
I've attached two instances of p1 of the Tribunal decision (follow the
links that Online groups adds to the email).
The first is the scanned page from MfE. The second is the scanned page
with OCRed text - now the text is selectable, copy & pasteable,
searchable. The second is what should be coming out of Government, if
non-native-digital documents are being distributed.
OCR is not 100% accurate. Any such solution needs to have a proof-reading
processing to correct errors.
On Sun, August 16, 2009 5:56 pm, Jim McLeod wrote:
> What you have here is the image of the actual genuine paper document
> signed
> by the Tribunal members. (You might even see their finger prints on the
> document ~8- [?]) ) That's why it is a PDF of an image. There are many
> opportunities to get the wrong version as multiple people edit these
> things
> both electronically and on paper. So it is good to get the signed one -
> each
> page initialled as well to show that it has been "read". You might even
> get
> last minute changes made in pen on the paper and they would also get their
> specific individual legal signature. This is the legal process and the
> legal
> document.
*sigh* I am aware that it is a legal document, and that a scanned image is
'potentially' the most likely to be 'valid'.
What I am complaining about is that a document comprising scanned images
is useless except for reading by a human, unless it also includes
accessible text (i.e. by OCR).
I've attached two instances of p1 of the Tribunal decision (follow the
links that Online groups adds to the email).
The first is the scanned page from MfE. The second is the scanned page
with OCRed text - now the text is selectable, copy & pasteable,
searchable. The second is what should be coming out of Government, if
non-native-digital documents are being distributed.
OCR is not 100% accurate. Any such solution needs to have a proof-reading
processing to correct errors.
Better access to public court records <https://www.recapthelaw.org/> via boingboingAlso a gem today from inside a government agency upon being probed about the use of tightly restrictive PDFs and imaged documents: "Information Management ... is all about protecting the integrity and content of the information received."
As you can see I am a bit slow about these things.
You suggest ATOM and RSS feeds. I know many agencies are using these to
supply "organisation communications", but I don't see how that will help
with the hearings process / public engagement process that triggered this
topic.
I am with you if you are seeking better access to information involved in
hearings. I have the same issue as you getting to grips with documents that
are practically inaccessible. The ideas about the OCR process for the PDFs
are useful. (In MfEs case at least some (all?) documents are available
online, in many cases this is not the case.) So yes, hearings documents
online (ideally in real time) in machine-readable format (XML?) would be
nice.
Another challenge was that hearings tend to be physical meetings. Hence a
big barrier is the requirement for physical presence at the meeting. So lets
ignore the "real time" situation for the moment.
Wouldn't a no.8 option, for after the event, be the use of mirror sites(?)
outside agency firewalls containing XML(?) documents for people to rummage
about in so that the agency can "protect the received information". This
mirror could be quickly established from a technical perspective rather than
waiting until the agency IT puts a portal in the firewall to view the
"protected info". Yes, not ideal but doesn't this release the information?
Doing the above jerry-jig will then help us on the way to the(?) goal of
citizen participation of getting the different ideas brought to the hearings
exposed and threaded.
What technology options are required/available for enabling and doing this -
the exposing of ideas?
On Tue, August 18, 2009 6:48 am, Jim McLeod wrote:
> You suggest ATOM and RSS feeds. I know many agencies are using these to
> supply "organisation communications", but I don't see how that will help
> with the hearings process / public engagement process that triggered this
> topic.
It would help because only 50% of the population read newspapers (and only
a small % of those probably scan public notices), hence they are not
notified of resource consent applications, public consultations etc. in
areas that affect them.
As soon as these public engagement opportunities are made available via
RSS/ATOM they can be parsed, searched, filtered, (Yahoo) piped, geo-coded,
mapped, calendared, etc.
For example, I am part of a national organisation (Whitewater NZ) that
want to be aware of every resource consent application for a hydro dam or
water extraction from a river. There's no current means I am aware of to
track such applications nationally. In my experience regional councils are
very poor and inconsistent at notifying affected parties such as kayakers.
One possible app I might look at during the Hackfest would be
screen-scraping the regional council consent application pages to generate
an Atom feed (unless others have better suggestions...).
On Tue, Aug 18, 2009 at 6:48 AM, Jim McLeod <email obscured>> wrote: > > What technology options are required/available for enabling and doing this > - > the exposing of ideas? Scribd [1], DocStoc and others provide document hosting - including parsing PDFs/word documents/excel files/etc, OCR, indexing them via search engine, making them available in different formats, and allowing embedding in other websites. I'm sure they'd have branded/ad-free offerings as well. In fact: "Strategies for how Scribd.com can be used by government organizations to help citizens find and view relevant documents" [2] Rob :) [1] http://www.scribd.com/ [2] http://forum.webcontent.gov/event/id/68513/Scribd.com-Uses-in-Government.htm
On 18/08/2009, at 7:54 AM, <email obscured> wrote:
> On Tue, August 18, 2009 6:48 am, Jim McLeod wrote:
>
> One possible app I might look at during the Hackfest would be
> screen-scraping the regional council consent application pages to
> generate
> an Atom feed (unless others have better suggestions...).
So Rob Coup and I want to work on a 'virtual API' for Government data
which would fit in with what you want. Basically the client queries
our API for some data and we go out to the various sources in whatever
format and return it in a standard format. Things like rates, consent
applications, etc would be a perfect fit for this. Looks like we could
have a couple of good use cases for building this at hackfest.
OKHearings and public notifications by RSS, Atom etc.
I missed that - sounds like a good idea.
As for my hearings documents, assume I now have them formatted and
accessible. (First step to heaven ~8-) )
Any tools for highlighting and assessing the issues (with evidence ie lets
get beyond opinion), the options and then the decisions across say fifty to
100 documents?
On Aug 18, 2009, at 8:45 AM, Jim McLeod wrote:
> OKHearings and public notifications by RSS, Atom etc.
> I missed that - sounds like a good idea.
>
> As for my hearings documents, assume I now have them formatted and
> accessible. (First step to heaven ~8-) )
For sharing events .ics should also be considered in addition to RSS/
Atom.
Jim McLeod wrote: > OKHearings and public notifications by RSS, Atom etc. > I missed that - sounds like a good idea. > > As for my hearings documents, assume I now have them formatted and > accessible. (First step to heaven ~8-) ) A few weeks ago I started http://consultations.org.nz which groups local council consultations by council and into past/current/upcoming. There are RSS feeds for each of these. The aim of the site is to make consultations available as html rather than PDFs, which most councils don't do, and allow discussion of the consultations, which most councils probably don't want to do. At the moment loading consultations is a slow manual process but going forward I hope to automate this as much as possible. There's still a loooooong way to go and I have a large to do list, but if I didn't start with something, no matter how basic, I'd have done nothing. I'd be interested in hearing your feedback.
Hi Ben,
Nice work! We will link to the site from open.org.nz when the new CMS
goes in.
In terms of the automation the HTML -> API project that Rob and I want
to hack on during the barcamp/hackfest might be able to help. If we
can write scrapers for each of the council sites it might at least get
them into your system automatically for some manual post processing.
On Wed, August 19, 2009 11:55 pm, ben wrote: > Jim McLeod wrote: > > OKHearings and public notifications by RSS, Atom etc. > > I missed that - sounds like a good idea. > > > > As for my hearings documents, assume I now have them formatted and > > accessible. (First step to heaven ~8-) ) > > A few weeks ago I started http://consultations.org.nz which groups local > council consultations by council and into past/current/upcoming. There > are RSS feeds for each of these. Great work Ben. > The aim of the site is to make consultations available as html rather > than PDFs, which most councils don't do, and allow discussion of the > consultations, which most councils probably don't want to do. HTML is best. My reason for discussing PDFs earlier is that tends to be how paper scans are published. > At the moment loading consultations is a slow manual process but going > forward I hope to automate this as much as possible. Are you doing any screen scraping or just manual loading? > There's still a loooooong way to go and I have a large to do list, but > if I didn't start with something, no matter how basic, I'd have done > nothing. I'd be interested in hearing your feedback. I'd be interested in assisting. I know my way around Drupal :-) I'm more interested in a similar service looking at notified resource consents. On the face of it, the two sites would share a lot of common modules, so perhaps we should discuss how to combine efforts. Have you considered using a service like Calais to autotag the consultations? http://drupal.org/project/opencalais
Sorry. Been offline for a bit.I think the scraping of Websites for feeding
to an API is a good step. It might encourage Councils and other agencies to
provide a more controlled, stable and formatted digital data feed along the
lines being requested.
The big motive here is that the agencies will likely want to avoid any
miscommunication about these events. Not good form to get public
notifications wrong.
The data from the "virtual API" can then be used to showcase innovative
demand/customer/citizen side products.
Yes - a very good idea.
Glen, Jonathan, thanks for the feedback and great suggestions. We can
chat more at barcamp/hackfest, I'm very keen to discuss scraping council
pages more with you, and also leverage your Drupal fu!
FYI all, Scraping of most council websites and publication of their documents is probably illegal. I've only checked Napier's, but read the following: "Portions of the Napier City Council information and material on this site, including data, pages, documents, online graphics and images are protected by copyright, unless specifically notified to the contrary. Externally sourced information or material is copyright to the respective provider. Use of website Images/Photos: Images/photos on this website can be used by students for school projects..." (Source http://www.napier.govt.nz/index.php?cid=struc/disclaimers&mid=802) It's important to remember that copyright protects every publication made. Proprietors have complete discretion over licencing for re-publication. I doubt fair use exclusions incorporate reproduction of entire articles. I would hate for the Police to come knocking at your door after you've been providing a valuable service Ben. The fact that you're not making a profit off of that information doesn't make it okay. If you're republishing others' content without written consent, I would seek some advice. Cheers, timClicks 2009/8/21 ben <email obscured>>
Yes - But not all councils have such restrictive copyright - http://www.horizons.govt.nz/default.aspx?pageid=340
Surely Napier CC are subject to Crown Copyright. They don't appear to be compliant with NZ Govt webstandards http://www.webstandards.govt.nz/copyright/ All it requires is for Ben to cite the source...
Totally agree, I used the word "most" in the previous email assuming that the majority would have a generic all rights reserved notice and hoping some would have less restrictive terms. A stocktake of licencing would probably take two or three hours. 2009/8/21 Glen Barnes <email obscured>> > Yes - But not all councils have such restrictive copyright - > http://www.horizons.govt.nz/default.aspx?pageid=340 > Even in that specific case, "[publication is okay, so long as] the material is not altered" can be interpreted very strictly: I just altered that copyright notice by adding a contextual clause at the start. Did I just break the law? My concern springs from a desire to maintain a good relationship with public sector agencies. The types of technical convergence that is being talked about on this list - allowing all govt data to be publicly accessible - will only happen with positive relationships and internal allies. If opengovt.org.nz is seen to be encouraging illegal behaviour by a potential opponent, maintaining those relationships becomes more delicate. Still - wonderful concept Ben. (Was working on something very similar locally to show around at Barcamp, looks like I've been trumped!)
Jonathan, Others will be better placed to answer this with more authority, but since I started this sub-thread... 2009/8/21 <email obscured>> > Surely Napier CC are subject to Crown Copyright. > They're not subject to Crown copyright, they create crown copyright. SSC's standards may have little sway in to local govt. I believe DIA's Local Govt Branch has more to do with keeping local authorities consistent. > > They don't appear to be compliant with NZ Govt webstandards > http://www.webstandards.govt.nz/copyright/ > Irrespective of any non-legislative guidance - every legal entity can set whatever standards it wants. My understanding is that the copyright owner's view of the world is what's important in the Copyright Act. Remember, the public sector is not one organisation. Even central govt agencies have inconsistent policies about all sorts of things. The core agencies (DPMC, SSC and Treasury) and Cabinet try very hard to provide a united front to the public. That doesn't mean life's like that in practice. Any non-legislative guideline can be ignored, interpreted or followed by each agency. Sometimes different branches/business units of agencies have their own processes.
On 21/08/2009, at 3:43 PM, Tim McNamara wrote: > Totally agree, I used the word "most" in the previous email assuming > that > the majority would have a generic all rights reserved notice and > hoping some > would have less restrictive terms. A stocktake of licencing would > probably > take two or three hours. > Or we could crowd-source it ;-) http://wiki.open.org.nz/List_of_Councils_and_Copyright_Notices
Here's an interesting case: Wellington City Council http://www.wellington.govt.nz/termsandconditions/ "Website visitors may reproduce, store and use the content of this website for personal, informational and non-commercial purposes only." On one level this seems fine. Now, would a bot/scraper be considered a visitor? I wouldn't think so, because a visitor to a website is generally regarded as someone using a web browser to access a website. If that's correct, then "Except as stated in the above paragraph, no portion of the content of the website, or the Council logo, may be copied or used without the written permission of the Council." 2009/8/21 Glen Barnes <email obscured>>
Wouldn't the best way be to actually approach the council in question,
tell them what you plan to do, why you want to do it, explain it is
non-commercial, and see if you can get explicit permission for your
use case?
These agreements are always up for discussion and negotiation ;)
It also serves a double purpose of making them aware of new and
interesting uses that citizens have for their data. If you just scrap
it, no-one is likely to find out about until they stumble upon the
mashed-up website.
Cheers Gav
Tim McNamara wrote: > They're not subject to Crown copyright, they create crown copyright. No, they don't. The materials they create *may* be subject to Crown Copyright. No-one creates copyright - it's not a thing. It is a legal condition that applies to things. There's also some debate over the applicability of Crown Copyright to local authority material as they are not, strictly speaking, part of the Crown. While they obviously form part of the State's structure, their legal status is reasonably independent. An example of this is that they have their own OIA legislation. > SSC's standards may have little sway in to local govt. I believe DIA's Local > Govt Branch has more to do with keeping local authorities consistent. Some of the e-government standards have had better uptake in the LG sector than among central government entities. Web standards is certainly one of these, though Cabinet was unable to direct local authorities to abide by them, and had to invite them, as they did for SOE's. LG Branch is more of a monitoring body than an authority. The local Government Commission, has some authority, but only over status and boundaries, not activities, policies or procedures. The Crown has some authority through the RMA and other legislation to demand that certain activities are undertaken, but each has to be bound by an act of Parliament, or an Order in Council. >> They don't appear to be compliant with NZ Govt webstandards >> http://www.webstandards.govt.nz/copyright/ They don't have to be. (See above) > Irrespective of any non-legislative guidance - every legal entity can set > whatever standards it wants. Um, no. Law is law. If a standard is not law, but is voluntary, then yes it is up to the entity to determine how much they meet this standard. Blanket statements of this nature are not helpful to discussion, Tim. > My understanding is that the copyright owner's > view of the world is what's important in the Copyright Act. Again, a blanket statement that is not correct. I suspect it is the phrasing that causes me concern, rather than your intent, but copyright is a legal state of existence and boundaries are firm. What we need to determine is where those boundaries lie. > Remember, the > public sector is not one organisation. Even central govt agencies have > inconsistent policies about all sorts of things. This is sadly very true, and one of the abiding disappointments to me when I was at SSC. > The core agencies (DPMC, SSC and Treasury) and Cabinet try very hard to > provide a united front to the public. That doesn't mean life's like that in > practice. Any non-legislative guideline can be ignored, interpreted or > followed by each agency. Sometimes different branches/business units of > agencies have their own processes. Usually through lack of though than deliberate policy. Some of those processes have been shown to be illegitimate when challenged. But this is generally true. The key point I want to make here is that it's as wrong to assume too much copyright as it is to assume that none exists at all. The best bet is to come to an arrangement with each council (there's only 85 of the buggers) which gives you an opportunity to talk about their data formats as well. Scrape what you need to show them a proof of concept and then discuss an ongoing relationship.
On 20/08/2009, at 9:15 PM, Tim McNamara wrote:
> I would hate for the Police to come knocking at your door after
> you've been providing a valuable service Ben.
I would love for the Police to come knocking at my door for doing
this. It would be an impressive feat of public-relations suicide for
a department to complain about someone publicising council
consultations.
I think the risk here is considerably lower than the amount of
discussion it has received. Yes, the risk might be greater for pure
data that they sell (e.g., scraping property information). Yes, do
tell them what you've done once it's done (forgiveness, not
permission). But consultations and notifications are fair game in my
eyes.
I agree with Nat.
In this case I have at least one Council willing to talk with reps from the
Group around these feeds.
In fact I will encourage them to join the discussion, though they may choose
to watch and listen rather than engage.
Go well tomorrow at the barcamp.
So lets get a proof of concept out at barcamp so we have something
that we can shop around and show people. Having something tangible and
not just a concept always seems help people 'get' why we are doing
these things.
On Aug 24, 2009, at 9:27 AM, Glen Barnes wrote: > So lets get a proof of concept out at barcamp so we have something > that we can shop around and show people. Having something tangible and > not just a concept always seems help people 'get' why we are doing > these things. > > Glen I have some ideas about how to put together a proof of concept application around an install of Kete (http://kete.net.nz) to pair data collected from witnesses with a official incident reporting data. I'll keep that in mind for the barcamp/hackfest.
I've had a long weekend in Aussie and am just catching up on email -
very interesting discussions while I've been away!
> If you're republishing others'
> content without written consent, I would seek some advice.
I specifically requested the larger documents I've had on the site from
the Wgtn CC and the Hutt CC as non-PDFs to make them easier to convert
to HTML, and told them what it was for, and in each case they were
extremely helpful.
Has/is anyone playing with http://code.google.com/p/ocropus/ As mentioned previously, I am interested in options for automated crunching 100s of documents (some as pdfs or even fax images) (think submissions) to flush out common themes.
On Fri, Sep 4, 2009 at 5:08 PM, Jim <email obscured>> wrote: > Has/is anyone playing with http://code.google.com/p/ocropus/ > > As mentioned previously, I am interested in options for automated crunching > 100s of documents (some as pdfs or even fax images) (think submissions) to > flush out common themes. We've recently been have some luck with http://code.google.com/p/tesseract-ocr/ to separate TIFFs with printed / typed Maori from TIFFs with handwritten English or Maori.
Hi Jonathan,How far did you get chasing this back to source? On Fri, Aug 14, 2009 at 4:13 PM, <email obscured>> wrote: > Hi all > > I just received the Special Tribunal report re a Water Conservation Order > on the Hurunui as a PDF comprised of scanned pages. > > I've immediately replied to MfE asking if they will have an accessible > text version available, and if so, when. > > <snip> > Would be nicer to fix at source... > > Regards > Jonathan > > http://huntdesign.co.nz > > I have been in "formal" contact with a "team leader" in the Environment Court Ministry of Justice | Tahu o te Ture who advice that: "*The electronic text of an Environment Court decision is only available in a protected and sealed document.*" In other words even though they word-process the document, they only make it available as an image. I did point out to the that this forces people to use manual or automated methods to recreate the electronic text of the decision, and that this introduces the risk of discrepancies as well as other inefficiences. This is hardly an add to productivity. The same intransigent response. Ideas on next steps? Lobby further up the chain?
On Mon, September 7, 2009 12:20 pm, Jim McLeod wrote:
> Hi Jonathan,How far did you get chasing this back to source?
Nowhere so far, not even a reply.
OkI think the source is the Ministry of Justice not MfE. Hence, MfE may have
some difficulty sorting out a response.
After a brief discussion with a colleague from the legal profession,
"protected" could mean without the possibility of annotations or early
drafts being attached (a la track changes) and in a form so that changes
can't be easily made, so the possibility of tampering is reduced. "Sealed"
refers to the application of the official court seal - in my case the
document has The Seal Of The Environment Court.
Of course a problem for anyone picking up a "protected and sealed document"
is that they can't take the document (presumably a copy) and easily compare
it against the "real" text. It is not easy to determine whether the copy has
been tampered with.
Still considering possible next steps...
Hi Jonathan,I have now discovered that the full (machine readable) text of
the Decision that I am working on is available in Brookers. Many other
Decisions are here also in full text machine readable form.
I wonder what the process is that creates these instances.
I will check for yours.
Jim McLeod wrote:
> Hi Jonathan,I have now discovered that the full (machine readable) text of
> the Decision that I am working on is available in Brookers. Many other
> Decisions are here also in full text machine readable form.
>
> I wonder what the process is that creates these instances.
>
From memory, Brookers get the image file and transcribe this. It's
historical process, as they did the same with paper to get it into their
publishing system.
Thanks. I wondered if that was the trick.
(In the early 1990s, I grabbed a paper copy of the RMA, sliced off the spine
and feed it through a batch scanner and OCRed it.
My CEO offered the electronic text to some other org's to help search the
act and pick up text. He quickly shelved the offer when some-one (I think
from Govt Print) mentioned it could violated government copyright.)
There has got to be a better way than the wagon at the bottom of the
cliff...
File a pro-forma OIA request for every OCR'd document and see what
happens.
Now that's something you could automate.
If I understand the problem correctly - that is how to get full text decisions of the Environment Court - then could I suggest the best approach for future decisions may be to lobby for the Environment Court decisions to be added to the existing public access site – http://jdo.justice.govt.nz/jdo/Introduction.jsp This would also enable it to be added to collections at the New Zealand Legal Information Institute – http://www.nzlii.org/ NZLI is the NZ branch of a worldwide effort to make statute, judgments and other key public legal documents available publicly. NZLI has judgments from most NZ courts and tribs, so why not the Env Court? I see that Brookers have the original image version of the Env. court decisions - with stamp of seal - and have presumably ocr-ed them to create an image/text pdf. I base this on the fact that some handwritten text on the image is, understandably, badly represented in the text of the pdf. BTW No copyright exists in judgments of any court or tribunal, whenever those works were made – Copyright Act 1994 s 27(1)(g)
> BTW No copyright exists in judgments of any court or tribunal,
> whenever those works were made – Copyright Act 1994 s 27(1)(g)
Snap. Was just about to follow-up with that...
Are you aware of any judgements relating to how the OIA impacts Crown
copyright? In theory Crown copyright persists, but in reality no sane
agency would try to enforce copyright over material delivered under
the OIA. But it would be nice to have a judgement or two stating that,
and the same for material released under LGOIMA.
Hi Jim I remember a similar thing way back where a NZ university was putting up online full text of legislation and they were obliged to take it down. What probably created the issue is not copyright but that the scanned/OCRd/reproduced text was presented, probably quite innocently, as just as good as the 'official' version. On the legislation.govt.nz site they talk about how they are gradually "officialising" the text of what is on the website, http://www.legislation.govt.nz/about.aspx#officiallegislation so there is a specific process running for guaranteeing the text (and format) is the same as the print version. Re: Brookers and similar, their legal editors add specific value with cross links and case law refs to the bills, acts and judgements. To clarify for the list members what Chris Esther said about Copyright Act 1994 s 27(1)(g), these are the works that by law, have no copyright. "No copyright exists in any of the following works, whenever those works were made: (a) Any Bill introduced into the House of Representatives: (b) Any Act as defined in section 4 of the Acts Interpretation Act 1924: (c) Any regulations: (d) Any bylaw as defined in section 2 of the Bylaws Act 1910: (e) The New Zealand Parliamentary Debates: (f) Reports of select committees laid before the House of Representatives: (g) Judgments of any court or tribunal: (h) Reports of Royal commissions, commissions of inquiry, ministerial inquiries, or statutory inquiries."
<email obscured> wrote:
> Hi Jim
> I remember a similar thing way back where a NZ university was putting up
online full text of legislation and they were obliged to take it down.
>
What actually happened was that a person at a university had downloaded
and was trying to mirroring Knowledge Basket's database, and was also
trying to do the same with Hansard. While the legislation was not
subject to copyright, the extra work that KB had done in presenting it
was assumed to be theirs. Never went to court, as the Clerk of the House
had words with the relevant chancellor and the site got pulled. Names
and dates would require a delve into my SSC email archive (shudder) but
it wasn't strictly speaking a university doing the mirroring, nor was it
about Crown Copyright.
It was between 2003 and 2005, I think, as that's when Russell and I were
having fun with the PCO ;-)
Ed
As you are aware IANAL but from my limited knowledge I would concur with you.
That there is no impact on Crown copyright by simply acquiring a
document under the OIA.
I have not encountered any cases on the subject though it would not
surprise me to see you as a party to such a case in the future ;)
On Mon, Sep 7, 2009 at 10:11 PM, <email obscured>> wrote: > Hi Jim > I remember a similar thing way back where a NZ university was putting up online full text of legislation and they were obliged to take it down. > > What probably created the issue is not copyright but that the scanned/OCRd/reproduced text was presented, probably quite innocently, as just as good as the 'official' version. On the legislation.govt.nz site they talk about how they are gradually "officialising" the text of what is on the website, http://www.legislation.govt.nz/about.aspx#officiallegislation so there is a specific process running for guaranteeing the text (and format) is the same as the print version. Re: Brookers and similar, their legal editors add specific value with cross links and case law refs to the bills, acts and judgements. > We have a whole lot of texts similarly covered Copyright Act 1994 s 27(1) (i.e. those listed on http://www.nzetc.org/tm/scholarly/tei-corpus-legalMaoriCourt.html , http://www.nzetc.org/tm/scholarly/tei-corpus-legalMaoriStatutory.html , http://www.nzetc.org/tm/scholarly/tei-corpus-legalMaoriCrown.html , etc), which we're mainly making available under a CC license. If someone wants them under a different license we're more than happy to work something out.
hi Jim Ocropus is now available in the newest version of Ubuntu, Karmic Koala. To get it running install the following: sudo apt-get install sudo apt-get install ocropus tesseract-ocr-eng The latter package is English language data without which Tesseract (and thus Ocropus) will not function. Then to run the character recognition against a PNG file do the following: ocroscript rec-tess test.png > test.html As an example I found this using google images - /_assets/http/www.nzqa.govt.nz/ncea/resources/chemistry/images/exp/90694-exp1-excell-rpt-7.jpg Parsing it through Ocropus I get the following reasonable output, http://pastie.org/708804