Posts Tagged ‘Public Data Corporation’
Dear Mr Cable
I read with interest yesterday your letter to the Prime Minister about some of the issues facing the UK in the future, and in particular the need for a vision and for a connected approach across government. This struck me as timely and useful, as it hopefully signalled the intention of a change in policy at one of the main roadblocks to innovation in improving government and fostering innovation.
I am referring to the policy of your own department – the Department for Business, Innovation and Skills – to restricting access to core reference datasets, such as the Ordnance Survey mapping data, postcodes, and company data, and thus not just stifling innovation and growth but preventing a consistent and connected approach across government.
Though much about the future is unclear one thing is certain, that we are increasingly living in a data world. In that world innovation – and democracy – depends on the ability to access and reuse data, particularly the core reference data on which other data is based: what area a postcode refers to, where something is located, who runs and owns the companies for which we work or which receive government money.
In fact, opening non-personal government data forms part of the government’s growth agenda, and it has already published a considerable amount. Yet much of this data is almost useless without the core reference to tie it together – data which is under the control of your department.
When I met with your then junior minister Ed Davey a couple of months ago on this subject, I asked him point blank whether the government was going to publish huge amounts of data under a licence which allowed free reuse, but was going to restrict access to the core datasets which tied these together, that were in fact the core infrastructure for our digital world? He said, ‘We’ve got some ideas for innovative charging models.’
Let’s put aside the fact that government departments aren’t the right people to come up with ‘innovative charging models’ – they don’t have the right skills, experience, and unlike entrepreneurs like myself they aren’t risking their personal money, but the nation’s future. Let’s focus instead on a ‘connected approach across government’. This would seem a perfect example of a relatively minor source of revenue (maybe as little as £50 million, according to the report published yesterday by Policy Exchange) preventing such an approach, and with it a route to how the UK will ‘earn our living in the future’.
In my own area, OpenCorporates has in a year grown to be the largest open database of corporate data in the world – without, I should add, any help, encouragement or cooperation from BIS. We have just released a new feature that allows search for directors across multiple jurisdictions, massively increasing the ability of journalists, fraud investigators, investors, civil society, customers and suppliers to understand companies. Needless to say, UK companies aren’t included in this list because this data is restricted to those who pay.
One vision for the future would include making the UK a genuinely open and transparent place to do business, for example making UK Companies House as open as that in New Zealand, where all data is available openly and without charge. It would include making the UK leaders in the field of open data, not just generating a world-leading ecosystem of companies such as we have in motorsport, but pioneering the use of open data by companies of all types and sizes. And it would include the government being able to reuse and publish its own data without the corrosive and restrictive licences placed upon it by the likes of Ordnance Survey, and thus have a truly connected approach.
You have it within your power to help enable that vision – I hope you will act on it.
Yesterday I received an email from a Cabinet Office civil servant in preparation for a workshop tomorrow about the Open Data in Growth Review, and in it I was asked to provide:
an estimation of the impact of Open Data generally, or a specific data set, on UK economic growth… an estimation of the economic impact of open data on your business (perhaps in terms of increase in turnover or number of new jobs created) of Open Data or a specific data set, and where possible the UK economy as a whole
How many Treasury economists can I borrow to help me answer these questions? Seriously.
Because that’s the point. Like the faux Public Data Corporation consultation that refuses to allow the issue of governance to be addressed, this feels very much like a stitch-up. Who, apart from economists, or those large companies and organisations who employ economists, has the skill, tools, or ability to answer questions like that.
And if I say, as an SME, that we may be employing 10 people in a year’s time, what will that count against Equifax, for example (who are also attending), who may say that their legacy business model (and staff) depends on restricting access to company data. If this view is allowed to prevail, we can kiss goodbye to the ‘more open, more fair and more prosperous‘ society the government says it wants.
So the question itself is clearly loaded, perhaps unintentionally (or perhaps not). Still, the question was asked, so here goes:
I’m going to address this in a somewhat reverse way (a sort of proof-by-contradiction). That is, rather than work out the difference between an open data world and a closed data one by estimating the increase from the current closed data world, I’m going to work out the costs to the UK incurred by having closed data.
Note that extensive use is made of Fermi estimates and backs of envelopes
- Increased costs to the UK of delays and frustrations. Twice this week I have waited around for more than 10 minutes for buses, time when I could have stayed in the coffee shop I was working in and carried on working on my laptop had I known when the next bus was coming.
Assuming I’m fairly unremarkable here and the situation happens to say 10 per cent of the UK’s working population through one form of transport or another, that means that there’s a loss of potential productivity of approx 0.04% (2390 minutes/2400 mins x 10%).
Similar factors apply to a whole number of other areas, closely tied to public sector data, from roadworks (not open data) to health information to education information (years after a test dump was published we still don’t have access to Edubase) – just examine a typical week and think of the number of times you were frustrated by something which linked to public information (strength of mobile signal?). So, assuming that the transport is a fairly significant 10% of the whole, and applying it to the UK $2.25 trillion GDP we get £9000 million. Not included: loss of activity due to stress, anger, knock-on effects (when I am late for a meeting I make attendees who are on time unproductive too), etc
- Knock-on cost of data to public sector and associated administration. Taking the Ordnance Survey as an example of a Shareholder Executive body, of its £114m in revenue (and roughly equivalent costs), £74m comes from the public sector and utilities.
Although there would seem to be a zero cost in paying money from one organisation to another, this ignores the public sector staff and administration costs involved in buying, managing and keeping separate this info, which could easily be 30% of these costs, say 22 million. In addition, it has had to run a sales and marketing operation costing probably 14% of its turnover (based on staff numbers), and presumably it costs money collecting, formatting data which is only wanted by the private sector, say 10% of its costs.
This leads to extra costs of £22m + £16m + £14m = £52 million or 45%. Extrapolating that over the Shareholder Executive turnover of £20 billion, and discounting by 50% (on the basis that it may not be representative) leads to additional costs of £4500 million. Not included: additional costs of margin paid on public sector data bought back from the private (i.e. part of the costs when public sector buys public-sector-based data from the private sector is the margin/costs associated with buying the public sector data).
- Significant decreases in exchange of information, and duplication of work within the public sector (not directly connected with purchase of public sector data). Let’s say that duplication, lack of communication, lack of data exchange increases the amount of work for the civil service by 0.5%. I have no idea of the total cost of the local & central govt civil service, but there’s apparently 450,000 of them, earning, costing say £60,000 each to employ, on the basis that a typical staff member costs twice their salary. That gives us an increased cost of £1350 million. Not included: cost of legal advice, solving licence chain problems, inability to perform its basic functions properly, etc.
- Increased fraud, corruption, poor regulation. This is a very difficult one to guess, as by definition much goes undetected. However, I’d say that many of the financial scandals of the past 10 years, from mis-selling to the FSA’s poor supervision of the finance industry had a fertile breeding ground in the closed data world in which we live (and just check out the FSA’s terms & conditions if you don’t believe me). Not to mention phoenix companies, one hand of government closing down companies that another is paying money to, and so on. You could probably justify any figure here, from £500 million to £50 billion. Why don’t we say a round billion. Not included: damage to society, trust, the civic realm
- Increased friction in the private sector world. Every time we need a list of addresses from a postcode, information about other companies, or any other public sector data that is routinely sold, we not only pay for it in the original cost, but for the markups on that original cost from all the actors in the chain. More than that, if the dataset is of a significant capital cost, it reduces the possible players in the market, and increases costs. This may or may not appear to increase GDP, but it does so in the same way that pollution does, and ultimately makes doing business in the UK more problematic and expensive. Difficult to put a cost on this, so I won’t.
- I’m also going to throw in a few billion to account for all the companies, applications and work that never get started because people are put off by the lack of information, high barriers to entry, or plain inaccessibility of the data (I’m here taking the lead from the planning reforms, which are partly justified on the basis that many planning applications are not made because of the hassle in doing them or because they would be refused, or otherwise blocked by the current system.)
What I haven’t included is reduced utilisation of resources (e.g empty buses, public sector buildings – the location of which can’t be released due to Ordnance Survey restrictions, etc), the poor incentives to invest in data skills in the public sector and in schools, the difficulty of SMEs understanding and breaking into new markets, and the inability of the Big Society to argue against entrenched interests on anything like and equal footing.
And this last point is crucial if localism is going to mean more rather than less power for the people.
So where does that leave us. A total of something like:
That, back of the envelope-wise, is what closed data is costing us, the loss through creating artificial scarcity by restricting public sector data to only those pay. Like narrowing an infinitely wide crossing to a small gate just so you can charge – hey, that’s an idea, why not put a toll booth on every bridge in London, that would raise some money – you can do it, but would that really be a good idea?
And for those who say the figures are bunk, that I’ve picked them out of the air, not understood the economics, or simply made mistakes in the maths – well, you’re probably right. If you want me to do better give me those Treasury economists, and the resources to use them, or accept that you’re only getting the voice of those that do, and not innovative SMEs, still less the Big Society.
Footnote: On a similar topic, but taking a slightly different tack is the ever excellent David Eaves on the economics of Toronto’s transport data. Well worth reading.
Update 15/10/2011: Removed line from 3rd para: “ (it’s also a concern that we’re actually the only company attending that’s consuming and publishing open data)” . In the event it turned out there were a couple other SMEs too working with open data day-to-day, but we were massively outnumbered by parts of government and companies whose existing models were to a large degree based on closed data. Despite this there wasn’t a single good word to be heard in favour of the Public Data Corporation, and many, many concerns that it was going down the wrong route entirely.
As I feared back when it was first announced, the proposed UK Public Data Corporation has got nothing to do with open data, and everything to do with protecting the interests of a few civil servants, turning back the open data clock to the dark ages of derived data and privileged access for the few.
However, the issue I’d like to focus on here, having last week attended a workshop on the PDC consultation is governance. [It's worth mentioning that I was the only one at the workshop without a stake in the existing public sector information structure, telling in itself.] And far from it being a dry, academic, wonkish subject, it is critical to the future of public data in the UK.
The reason this is so contentious is twofold:
- The consultation on the PDC has been drawn very narrowly, trying to get respondants to choose between a set of options that are all bad for open data, and ultimately democracy. “So, open data, would you like a bullet to the back of the head, or to be slowly drained of blood?”
- There are clear conflicts of interest between the wider interests of society, and those of the Shareholder Executive – the trading funds such as the Ordnance Survey and Land Registry who are the very roadblock that open data is supposed to clear, but yet who crucially seem to be driving the PDC.
Now, from their perspective, I can see the appeal of keeping everything cosy and tight, particularly if there’s a chance the organisations being floated off, and with it considerable personal enrichment. But public policy shouldn’t be driven by the personal interests of civil servants, but what is in the interests of society as a whole.
In fact, the governance of the Public Data Corporation, and the rules by which it operates were the one thing that everyone at the workshop I attended agreed upon. In fact more than that, it was agreed that the delivery of its duties should be separate both from the principles by which it operates (which should be for the benefit of society) and the independent body that needs to ensure it sticks to those principles.
But here’s the kicker, the Transition Board for the PDC (which will oversee its membership, structure and governance) is, I understand, meeting on October 25, two days before the consultation ends.
When I asked this meeting, and whether the consultation was a done deal, I was told, “The governance of the PDC is not being consulted on.”
This is both rather shocking, and shameful, and for me means there’s only one viable option if the UK is serious about open data: to send the whole PDC concept back to the drawing board, and this time to come up with a solution that is focused not on civil servants’ narrow personal interests, but on building a ‘more open, more fair and more prosperous‘ society (to quote the Chancellor).
A couple of days ago, there was a brief announcement from the UK Government of plans for a new Public Data Corporation, which would “bring together Government bodies and data into one organisation”.
A good thing, no? Well, up to a point, Lord Copper.
I tweeted after the announcement: “Is it just me, or does the tone of the Public Data Corp make any other #opendata types uneasy?” From the responses, I clearly wasn’t the only one, and in my discussions since then it’s clear there’s a lot of nervousness out there.
So, what is it, and should we be afraid? The answers are ‘Nobody knows’, and ‘Yes’.
To flesh that out a bit, none of the open data activists and developers that I’ve spoken to knows what it is, or what the real motivation is, and remember these are the people who did much to get us into a place where the UK government has declared that the public has a ‘Right To Data’ and that the excellent ‘Open Government Licence‘ should be the default licence.
In that context, the announcement of a ‘Public Data Corporation’ should be be treated with some wariness.
However, this wariness turns into suspicion, when you read the press release.
First the announcement is a joint one from the Cabinet Office minister Francis Maude (who seems to very much get the need for open public data in the changed world in which we live) and from Business Minister Edward Davey, who I know nothing about, but his department BIS (Dept of Business, Innovation & Skills) has very much not been pushing for open data, and in fact has in the past refused to make data it oversees openly available.
(My sources tell me the proposal in fact originated from BIS, and thus could be seen as an attempt by the incumbents to co-opt the open data agenda, as a way of shutting it down, smothering it if you like.)
Second, despite the upbeat headline “Public Data Corporation to free up public data and drive innovation” (Shock horror: org states its aim is to innovate & be successful), the text contains a number of worrying statements:
- “By bringing valuable Government data together, governed by a consistent set of principles around data collection, maintenance, production and charging[my emphasis], the Government can share best practice, drive efficiencies and create innovative public services for citizens and businesses. The Public Data Corporation will also provide real value for the taxpayer.“
The idea of ‘value for the taxpayer’ is the same old stuff that got us into the unholy mess of trading funds, and the gordian knot of the Ordnance Survey licence wich is still being unpicked. This nearly always translates as value we can measure in £s, which in turn means what income we’ve got coming in (even if it’s from other public sector bodies).
- “It will provide stability and certainty for businesses and entrepreneurs, attracting the investment these operations need to maintain their capabilities and drive growth in the economy” – quote from Edward Davey.
If I were a cynic I’d say stability and certainty translates to stagnation and rent-seeking businesses, which may be music to civil servants’ ears but does nothing to help innovation. We’re in a rapidly changing world. Get over it.
- “bringing valuable Government data together, governed by a consistent set of principles around data collection, maintenance, production and charging”.
If this is the PDC’s mandate I think it could end up focused on the last of these, short-sighted though that would be.
- “It will also provide opportunities for private investment in the corporation.”
Great. A conflicting priority, to delight the bureaucrats and muddy the focus. Keep it small, keep it simple, keep it agile.
Finally, there’s no mention of open data, no mention of the Open Government Licence, the Transparency Board and only one mention of transparency, and that’s in Francis Maude’s quote.
If you’re a natural cynic, you’ll just say the government has already decided to flog everything off to the highest bidder. If you adopt that position, and give up without a fight, the people in Whitehall and the trading funds who want to do that will almost certainly win.
However, if you believe me when I say things are finely balanced, that either side could win, and enough well-organised external pressure could really make a difference over the next year, then you won’t just bitch, you’ll get stuck in.
He’s not wrong there. We’ve got perhaps 6 months to make this story turn out good for open data, and good for the wider community, and I suspect that means some messy battles along the way, forcing government to take the right path rather than slide into its bad old habits, perhaps with some key datasets, which should undoubtedly be public open data, but are currently under a restrictive licence.
I’ve got a couple in my sights. Watch this space.