Improving access to non-domestic energy consumption data

I recently wrote a post describing the data ecosystem for non-domestic energy consumption data in the UK. In that post I summarised my current understanding of the different actors involved in that data ecosystem, and some of the challenges of trying to access data from the perspective of a third-party service provider. In an earlier … Continue reading Improving access to non-domestic energy consumption data

How could watermarking AI help build trust?

I've been reading about different approaches to watermarking AI and the datasets used to train them. This seems to be an active area of research within the machine learning community. But, of the papers I've looked at so far, there hasn't been much discussion of how these techniques might be applied and what groundwork needs … Continue reading How could watermarking AI help build trust?

24 different tabular formats for half-hourly energy data

A couple of months ago I wrote a post that provided some background on the data we use in Energy Sparks. The largest data source comes from gas and electricity meters (consumption) and solar panels (generation). While we're integrating with APIs that allow us to access data from smart meters, for the foreseeable future most … Continue reading 24 different tabular formats for half-hourly energy data

12 ways to improve the GDS guidance on reference data publishing

GDS have published some guidance about publishing reference data for reuse across government. I've had a read and it contains a good set of recommendations. But some of them could be clearer. And I feel like some important areas aren't covered. So I thought I'd write this post to capture my feedback. Like the original … Continue reading 12 ways to improve the GDS guidance on reference data publishing

Brief review of revisions and corrections policies for official statistics

In my earlier post on the importance of tracking updates to datasets I noted that the UK Statistics Authority Code of Practice includes a requirement that publishers of official statistics must publish a policy that describes their approach to revisions and corrections. See 3.9 in T3: Orderly Release, which states: "Scheduled revisions or unscheduled corrections to … Continue reading Brief review of revisions and corrections policies for official statistics

The importance of tracking dataset retractions and updates

There are lots of recent examples of researchers collecting and releasing datasets which end up raising serious ethical and legal concerns. The IBM facial recognition dataset being just one example that springs to mind. I read an interesting post exploring how facial recognition datasets are being widely used despite being taken down due to ethical … Continue reading The importance of tracking dataset retractions and updates