All posts in the topic Paper on Managing Open Data (Short link)
Summary
- There are 13 posts — by 9 authors — in this topic.
- Latest post made by Jonathan Gray at May 12 03:00 NZST
I am writing a paper titled "Managing Open Data", which will be published in a
*book* later this year (US publisher).
I know it is a bit oxymoronic to capture in dead trees a topic that is so
continuous beta. But I am going to give it a go.
I would welcome any suggestions on key points that you think should be
included.
The biggest barrier I'm facing to accessing government data is
information released only in PDF format.
Laurence, please state that tabular data should be released in formats
other than PDF. Either CSV or HTML is much easier to machine process.
Rob
On 6 May 2010 12:21, <email obscured>> wrote:
> I am writing a paper titled "Managing Open Data", which will be published in
a *book* later this year (US publisher).
>
> I know it is a bit oxymoronic to capture in dead trees a topic that is so
continuous beta. But I am going to give it a go.
>
> I would welcome any suggestions on key points that you think should be
included.
I agree with Rob.
I use to fight paper now I fight PDFs.
And the volume is orders of magnitude greater - people had to pay to deliver
paper, now the transmit cost of the information has gone.
But most people still end up printing the PDF to read them, as they come
locked as A4, and are usually completely intractable (no text select, no
search ,,,)
On 2010-05-07, at 05:57 , Jim McLeod wrote: > But most people still end up printing the PDF to read them, as they come > locked as A4, and are usually completely intractable (no text select, no > search ,,,) Slightly OT, and not sure how many people are aware of Walking Papers, but it is a pretty cool use of printed pdfs. One of the few ;) <http://walking-papers.org/> Cheers Gav
From:
"Brent Wood" <email obscured>>
Add sender to Contacts
To:
<email obscured>
Cc:
<email obscured>
Hi Laurence,
I suggest you mention the three core requirements for data to be both open &
useful.
There is a trend in NZ & elsewhere to seize the CC by-sa licence with a sigh of
relief & say problem solved. Unfortunately this is only part of the problem, as
discussed in many forums & in the licences that recognise the differences
between data & creative works.
Licence: CC by-sa is pretty reasonable.
Attribution: Any attribution requirement must be flexible enough that mashups
with other data are permitted. Some attribution requirements, even under CC
by-sa, require attribution of data & derivative works that are incompatible
with the attribution requirements of other datasets. Open licence, but still
not able to be used freely in mashups/remixes.
Provenance/metadata: If I have 3 street map datasets, for example, for these to
be useful I need to know where they originated (credibility), how they were
generated (reliability/accuracy/precision), when they were generated & from
what underlying data sources (relevance, most current), etc.
Any datasets released without all three aspects covered generally might as well
not be released at all, except for very limited use.
Can you post a link to where your paper is published?
Thanks,
Brent Wood
Hi Laurence/all Does the concept of Open Data continuous beta mean that the Information Management issues are new? When I was with the Ministry of Fisheries, we implemented the Policy Framework for Government Held Information around 2000, as part of our Information Management Strategic Plan. It seems to me, that the issues we overcame then, around metadata / context/ usage, are the same issues now with Open Data. Edwin Bruce as MFish CIO back then, would be a valuable person to interview. When I was with the e-government unit, we did a lot of pioneer work in Internet concepts around delivery that are still relevant: - microformats; the concept that web pages are well structured data pages, accessible by people and machines; was an early experiment in machine readable semantics. There are some good lessons there, about why its hard to implement. - real-time archive; the concept that a reliable Government Shared Network could enable a real-time authoritative data source with reduced IM costs (duplication, management overheads, etc) http://blog.e.govt.nz/index.php/2008/11/10/nz-government-information-revolution/ - aggregated directories from distributed sources; the realisation that agencies have no interest in maintaining copies of data in other directories, but if we can semantically markup their data, we can virtually aggregate it using search engines http://research.elabs.govt.nz/new-zealand-government-feed-standard-2009/ Another major concept that I would like to see develop, is around supply chain open data. This concept is only feasible with our development in networks/bandwidth/standards etc over the last 10 years. Rod Drury gave an example of the value here: http://www.stuff.co.nz/business/3519667/1b-from-milking-the-cloud Personally, I don't believe any one party should hold the information, but all parties should create a trust that has its own governance. Would it be semi-open data? There's national competitive advantage in our stakeholders having access, but not our competitors. How do you resolve the stakeholder tensions? For example (barcodes), the suppliers want to publish their barcode/product information (open data), but are reluctant to support customer reviews attached to their products; and their retailers are unhappy about allowing pricing information to be included. Government might want to attach its certification information, but its not funded for that. The customer stakeholder is vitally interested in all those other aspects. As a final chapter in your book, I think it would be useful to talk about personal data. There is value in sharing parts of my personal information with groups of people (doctors, teachers, family). Lawrence Lessig said to me once that it would be reasonably easy to implement a parallel creative commons privacy framework; think about how we could then leverage this value? Mike Pearson http://www.linkedin.com/in/mikepearsonnz On 6/05/2010 11:21 p.m., <email obscured> wrote: > I am writing a paper titled "Managing Open Data", which will be published in a *book* later this year (US publisher). > > I know it is a bit oxymoronic to capture in dead trees a topic that is so continuous beta. But I am going to give it a go. > > I would welcome any suggestions on key points that you think should be included.
Thanks everyone for the ideas so far.
The pdf issue is one that I wil definitely cover (I was thinking of calling the
section "no-one gets out of here alive").
The provenance and integrity of open data suggest that managing open data is an
eco-system challenge - how to design open data to be self managing - rather
than a task or responsibility.
Just a clarification Mike, I am doing one chapter in a book on Managing
Government Information - a whole book is more than I want to commit to at this
stage. I will be able to distribute my chapter once the book is published.
More ideas....
This paper indicates the fork in the road - is the semantic web with richly linked data, each carrying it's own metadata, really viable, or just a theoretical dream. http://pewinternet.org/Reports/2010/Semantic-Web/Overview.aspx?r=1
On 7/05/2010, at 7:30 PM, <email obscured> wrote:
> The provenance and integrity of open data suggest that managing open
> data is an eco-system challenge - how to design open data to be self
> managing - rather than a task or responsibility.
Actually, I think it'll always be someone's task or responsibility if
we're talking about government data. Government is phobic about
having "bad data" out there--an incorrect datum that might cause
someone to die, get lost, or make a poor investment. They'll always
want to have the dataset be *someone's* job within government, not
wanting to put the govt imprimatur on unmanaged third-party data.
That isn't to say that there can't be active third-party
contributions, just that it's someone's job to go through and filter
those and ensure they're up to snuff.
just a theoretical dream
In practice, to my knowledge, conversion of data to RDF requires a
costly waterfall process. I'd rather government provide unique
persistent URIs for subjects of interest, alongside the raw data.
Leave communities and the market to decide if there's value in
converting that to RDF. In my experience, the service you want to
provide to people doesn't need the RDF step.
Rob
On 7 May 2010 08:36, <email obscured>> wrote:
> This paper indicates the fork in the road - is the semantic web with richly
linked data, each carrying it's own metadata, really viable, or just a
theoretical dream.
If Gov't depts, or any corporate are to release data, then the No. 1 beauraucratic rule of "cover thy butt" will kick in. It will need to be done via guidelines & standards across a very large playing field with lots of players. I have also seen ridiculously incompetent use made by supposedly professional people of data that was not collected or intended for the purpose it was used for. All of which means someone in every organisation releasing data in such an environment must be seen to do it responsibly, ensuring privacy/confidentiality constraints are not broken, clear fitness for purpose metadata is needed, along with as many waivers of liability as can be squeezed in. This is required at the point of release, not afterwards. Maybe overstating it somewhat, but to some extent the eco-system challenge is in ensuring that inappropriate use of open data does not result in too much mayhem. Open data means lots of statistics: lies, damn lies & statistics :-) See: /_assets/http/intranet.iucn.org/webfiles/doc/SpeciesProg/RL_Guidelines_Data_Use.pdf for an example of fitness for purpose metadata. Brent Wood
This might be useful, which looks at various legal and technical aspects: http://www.unlockingaid.info/ We're also starting work on a new Open Data Manual based on this and other work: http://lists.okfn.org/pipermail/okfn-discuss/2010-April/007244.html And we are going to cover similar ground in a forthcoming report on Open Government Data: http://opengovernmentdata.org/ And we've also got rawdatanow.com on the back burner, which we intend to use for examples, stories and case studies around getting hold of data in raw, machine-readable form. This will be aimed at the general public, government departments who publish data, and so forth! All the best,