Citizen Science with Python
Talk 30 min
Transcript
Read transcript
right so Who am I my name is Ian Oswald I've met some of you yesterday it's very nice to be here this is my first time traveling to Lithuania and I'm very happy to be invited by IDUs and the organising crew so I'm an interim chief data scientist this is a job title that I made up a couple of years ago I've been creating my own career for 15 years now and I keep I keep trying to figure out what it is that I'm doing and why am i doing it and so I keep coming up with new job titles and one of the problems in our field is rampant job title inflation so people have one year of experience and they try to claim their a senior person in data science or machine learning or whatever because those jobs pay better and when I realized that I can make up my own job title because I ran my own company I decide that I'm chief from clearly chief I'm chief data scientist and the t-boy I'm everything but rather than claiming any job title in between I thought well I I can't be any bigger than the chief so that's who I became so I've been working in the field of data science and artificial intelligence for almost 20 years I did my master's degree in 1998 I nearly did a PhD and then I was tempted into the world of startups and I worked in a start-up on the University campus in Sussex University in the UK down on the south coast and I kept meaning to go back and do a PhD but it turns out once you started earning some money it's quite nice to keep Herning some money and then you buy a house and keep earning money and in 20 years past which is which also means that unlike many data scientists who who graduate and had a PhD and new math and knew what they were doing some of us just make it up as we go along so I've had 20 years of making it up as I go along and I'm pretty proud of that because it turns out you have to be pragmatic and just just figure out what's required in the environment around you I'm also an author of O'Reilly high performance Python I co-authored that with Anna can chap we should Gorelick five years ago we've just started working on the second edition to be published next year O'Reilly are very kind to their authors and they send boxes of books around for signing at conferences so I have 20 copies of this book here to give away today and we'll do that in the first coffee break the volunteers out the front will do some kind of organization so you don't need to leave a talk to come down for it just during the coffee break once the coffee has start it will do the signing there 20 copies free to give away I'll sign it for you so you just have to come and queue up when the when the volunteers say that the queue is starting I've also started to do video tutorials with plural sites to anyone no plural site there's that Brad oh yeah oh wow lots of you owe more than in my London user group that's that's amazing so I've got a video on the site published about three weeks ago now introducing data science for engineers who are getting into the field and want to visualize their data which is the theme of today's talk and the thing that I'm most proud of is the PI data organization I'm one of the cofounders of PI data London our London meetup he's one of a hundred and forty and remember that number had come importantly to run one of a hundred and forty meetups around the world that have grown up in the last seven or so years we started six years ago we were the first outside of the US and we've now grown to the largest PI data on the planet and the largest data science open open-source data science group on the planet for Python with almost 10,000 members so we've got a huge membership in London 200 people turning up every month it's crazy and really good so I'm really proud of this community that we built and I'm really really happy to see the community building up out here and the PI data tracked you've got evolving here and the other PI data organizers we have here so today I'm gonna tell some short stories on citizen science with Python so doing good with our pythonic and programming skills using data I'm going to provide some tips for how you could do the same and then there'll be a short live demo at the end to show some of the tools that we can use to do this before I start I want to make sure I hope I'm pitching this to the right audience how many of you identify as data scientists or wedding put your hand up okay and then put your hand up if you identify as software engineers yeah okay good right you're you're my main audience data scientist sorry you're gonna have seen some of this before the stories hopefully will be new to you but the demo won't be so the first story I'll tell is about the Macedonian air quality mmm so this was a oh now North Macedonia I'm reminded when I heard this story it was still Macedonia under the negotiation with Greece that was renamed in North Macedonia a year or so two years ago so I heard this story at PI date Amsterdam so one of the con one of the PI data conferences held in Amsterdam by chap called Gordon and he was telling it when he's a bit older but he was telling the story of when he was 21 years old this is a few years back and the the air quality situation hasn't really changed unfortunately so this story is still incredibly relevant so Goergen started programming when he was 21 and then this investigation comes out of that and separate two that he was aware he had some awareness like everyone that the air quality wasn't very good and this picture is not cloud that has come down into the city this is pollution that has risen up and can't escape the city and so every summer city fills up with pollution and you can see if you fly you can see the tops of the buildings but if you're on the ground you can't see anything and this is normal this is known as the smelly fog and it's a normal situation no one is worried about this because hey it happens every year right it's just normal the government says that it's okay Gordon got some data and was trying to figure out what the data was telling him and he couldn't understand that he was programming how he couldn't normalize his data it didn't seem to make sense and the more he normalized it the more he couldn't understand why he is a junior programmer on his first good project just couldn't make these numbers make sense and it turns out the reason was that the the data showed that the air pollution was 20 times the EU pollution limits four times Beijing's pollution we know the Beijing's pollution is really bad so this is crazy bad and Gordon's sitting there saying but my numbers they're so big it can't be possible this is ridiculous it turns out now it's the truth that's the air that he's sitting him when he's coding away it was really really awful and so this turned into a bit of a political movement and so the interesting bit of this story is this is someone who's learning to program ends up turning into a bit of a political activist so initially Gordon got a Jason data dump and I asked at the end of his talk why why did you get a JSON data dump if the day that if the data was so bad for the pollution why the heck would the government published the data they set out one of the requirements of us coming into the EU is that the EU gave us some things and some of the things they gave us were air pollution monitoring devices but a requirement was we had to publish the data as open data so the data was published but it was undocumented it wasn't meant for anyone to actually read it and do anything with it and there was certainly no push from the government to highlight just how bad the data was so gorgeous data and then learned to program with it realize just how bad the situation was he made a website first of all and then that became a mobile app had it turned into a proper app over some time and but as he tells the story a million people ended up being engaged in this when he got some publicity within a month and then the news media started to run stories with it and people started to realize that the smelly fog this standard situation was not normal and it was not a good thing the visualizations got printed out that's a member of parliament from the right hand side showing out showing pictures in Parliament from the app and this led to intense discussion from the opposition party in government saying this is ridiculous down with the main party you should vote us in we'll fix everything and it got to the point that a minister in the government challenged Gordon that his app was in error that data was wrong it was all a lie and it was bad so Gordon counter challenged back to the Minister publicly and if you can demonstrate that the data is wrong I will take everything down however if you challenge publicly and it turns out the data is right you will resign unless he tells it the minister then didn't come back so they've got a political movement going at this point he's found other people they're getting they're doing lots of publicity they've got support in Parliament and it ended up driving government policy so it ended up becoming part of the election manifestos for both parties to clean up the air in in the next election cycle which is a pretty amazing outcome now I did also follow up on what happened after that and it turns out there was a change in government the policy was changed and then things kind of slowed down again the things kind of went back to business as usual but the population are much more aware now there are more pollution monitors around and people are more used to asking questions about this and challenging the government and so there are many more people citizen activists using the data that comes out to popularize the idea that hey bad air kills people and you definitely kills we've got a court case in the UK in London at the moment with a person who died and there's a court case that determine if it can be if pollution can be determined to be the main cause of her death and then that will possibly lead to suing the government afterwards so that's a pretty big thing of happening in the UK and this is happening throughout Europe and around the world now so it's good we finally figured out the bad air kills people get was the plants kills the environment and it's bad and we shouldn't do it and so I loved their development of this story a future piece that came out of this is they're now collaborating with the EU Space Agency to get satellite data to look for hot spots or pollution hotspots things that are messing up the air locally because it turns out they're a bunch of illegal polluters including an ex British waste pollution plant that many years ago was sold to Macedonia because it was illegal to run in the UK due to pollution limits and it wasn't illegal in Macedonia and so these things get highlighted and then fixed so I really like this so how would you start well get a public data set there's lots of public data in the EU you're in the EU so you can get some of this public data you can get your own data set of course from your own company your own academic institution but there's lots of public data load it and invest negated an important thing to realize is most people most of the population of this country most of the population of the EU cannot access this data they don't know what a CSV file is or what a JSON decode that looks like they don't know how to draw a graph and so if you do something and then you show it to someone who is an activist or an environmentalist somebody who's interested who doesn't have programming skills you will have an amazing collaboration in no time simply because you've got the keyboard skills so you make some graphs use matplotlib in Python or plotly as a public plotting tool and then you tell a story and then let me say that again this is the important bit get the data figure out what's in it and tell a story and by telling a story you will get other people to work with you and begin to cause change around you so here's the second story updating outdated medical results this is a really nice one so I heard about this it was actually passed around the PI data organizers after Anna and I'm going to butcher her name nicely I can't butcher her name Anna Anna s anyone told me how I would pronounce that name in Polish Strieber stiva even that I'm gonna butcher an arrest anyway I do not wish to make a name wrong so Anna gave her a five minute lightning talk just like the lightning talks that we had here last night and we'll have tonight's talking about Friedman's curve who's familiar with Friedman's curve put your hand up it's around childbirth one hand two hands okay so the frequent curve developed in 1955 is around childbirth so the curve was developed by measuring for a number of patients a number of women going through childbirth by Friedman and his team back in 55 about cervical dilation during the hours of birthing so when the baby's going to come out the cervix has to dilate it takes about 6 to 12 hours and then when it's the right size the baby pops out birth done everyone's happy does it happen in a consistent way well Freedman general gathered the data to tell the first story and provide the first medical advice and it has become the standard medical advice around the world and this medical advice says there is one standard way that the baby comes out due to cervical dilation only it turns out there is not just one way that this works so there a set of stages of labor the latent phase takes about eight hours and then there's an accelerated phase cervix dilates quickly at this point and then things happen quite fast if you're a midwife and you're monitoring one of your patients if the cervix doesn't dilate on schedule you have to make a call well maybe there's a problem and if there's a problem maybe we have to intervene so maybe we introduce drugs to make the cervix dilates and the birthing process speed up to fit the curve because the curve is normal or maybe we need to go and do more maybe we need to have a cesarean operation or maybe we have to do some kind of other intervention but these are pretty big medical steps they will have a negative impact upon both the mother and the child and they're costly for the hospital as well but we know we have to do these things because the data tells us from 1955 that this was the case but it turns out sixty years have passed and lots of things have changed particularly back then they had different drugs that they used they often provided drugs to all mothers at the beginning to change the birthing process they used forceps during delivery at least some of the time that's generally not used anymore and the women were of a different age and a different kind of health in general so 60 years of education in the birthing process has changed the set up entirely but we still follow this graph from 6070 years ago and so Anna ran her colleagues gathered new data to update these outdated results what we see here is a so the red dotted line is the Friedmann curve and then we had the bar charts their time along the bottom so at one hour for a number of subjects that were being measured against this bar charts and the cervical dilation was between zero and three centimeters and then by say six hours most of the patients had a cervical dilation of five and ten centimeters but those whiskers they're the long bars suggest there are some outliers so it's still the case that at five or six hours some people haven't really started their dilation process yet but most of the women are in those bar charts above the red lines that most of them have their process occurring faster than the free bun curve would suggest then we have that rapid change at around nine hours and those dots those are individual instances of women who are still not going through the process as fast as everyone else everyone else has already gone to 10 centimeters they've had their baby and the process is over so this tells us that the the classic curve doesn't actually fit one real-life medical unit and actually this is one of many studies around the world and they all show that the birthing process around the world does not match the standard Friedman curve so what do we do with this first of all this highlights the problem then they they're a tiny bit of machine learning which is really sensible they made a single decision tree so for the data scientist in the room she's not super advanced this is not deep learning this is a simple decision tree a set of if-then rules the top rule says are you a mother who is having her first birth or second or subsequent births it turns out the first birth takes longer and then birth to three four tend to happen quicker if it's your first birth you go down the left-hand side of the tree and then if you are under a certain weight you go further down the left and if you are over a certain weight you go down the right that will affect the bursting time and then age becomes a fracture and a few other factors come in the really nice thing about this is they could sit down with midwives who have no programming skills at all no data analysis skills and work through some of these answers the Midwife would say yeah sure yeah that completely makes sense I see that and of course they see that there's the data they're gathering from their birthing unit and so they're using this to begin to change the policy of how they deal with patients and particularly how they make a decision about whether they need to go and do a bigger medical intervention or whether actually although the mother is not on the Friedmann curve the mother is inside the expectations for the data they've gathered and so there are no problems occurring so how would you start and this can be this could be in medical science but this could be in your company this could be in your academic institution the to be with government-provided open-data so check for outdated assumptions it turns out the world moves fairly quickly and if you're using things from 70 years ago maybe they're not true anymore in your business things probably changed in the last five years if you've changed your customers and how you're making your products things have probably changed and yet probably people are still using the old methodology in the old date of the old assumptions the old decision points and maybe they're wrong maybe you can have quite a big impact by just questioning those old assumptions and getting a bit of data so again you gather your data and you visualize it make simple models the number of people in data science who rush to make the next most complex model it has to be tensorflow and i need to use apache spark and know almost always all you need to do is get some data and put it on a USB stick analyze it draw a couple of graphs and then put a few if-then-else statements in there and that will actually let you make the first big new decisions in a way that's entirely understandable and debuggable and that all of your colleagues can communicate with you around and then yeah you can build up iteratively as things improve so now where we've done a bit of personal house we've done a bit of medical health what about environmentalism so a colleague of mine göransson back in London super smart chap he runs the London or he was at one of the co-founders at the London machine learning meetup he also works with a startup called Oxbow tikka they do self-driving vehicles up in Oxford and he's had a long had an interest in autonomous vehicles and it turns out prior to his work on autonomous cars he was working with the charity trying to save orangutangs and this is hands and ear I think so down in the middle of Africa and here's the challenge you have you have a Ranga tangs look could be sold as pets and so caged baby orangutangs and then they get abandoned at a certain point in their life when they get big or due to deforestation and farming we end up with displacement and so volunteers have to go in and rescue the orangutan when I cost money of course these are financially driven organizations they're typically charities so these orangutangs are rescued and then when they're relocated and you release them you want to make sure they're doing well have we done a good job if we just put them somewhere else where they died and there will be a bad outcome we want to make sure we put them somewhere where they live and they live a healthy rescued happy life and so they have to be relocated and tracked how do you track an orangutan well it doesn't have a neck like we do it has very fat neck so you can't put a ring on it you can't give it a watch and you can't attach anything to it it's a very intelligent animal if there's not one things being attached to it and so what you do is you take a little micro transmitters and you put them under the skin so you cut the skin open pop one of these inside then it lasts for several years it turns out they're reasonably reliable has a set of blog entries on this and they're they're not actually super reliable they tend to break quite quickly because they're small they go under the skin their range is not very long and can someone tell me if you're in the rain forest what happens with radio waves and the rain forest Segan muting wood you mean by that yes the range of signal is decrease so the water interacts with radio waves and when you've got thick canopies and lots of mist your radio range is significantly reduced when your device is tiny and you're the thing you're tracking is intelligent and very mobile it's quite hard to track them so the way you track them is and the the animal is released and then a tracker a human has a radio beacon has a range of 30 maybe 50 meters they go to where they release the animal yesterday and they see if it's still there today and if it isn't they go walking and after a while maybe the device starts beeping it's not directional so they just have to walk backwards and forwards until it beeps louder and then they home in and they find the orangutan brilliant they sleep the next day they wake up they listen to the device it's not beeping they start to walk again and they go to try to find the orangutan so they check for days and then weeks and then months to find the DISA Ranga tang or other orangutan goes and see how they're developing but that's a very manual process the question can we use drones on the right hand side there to fly around and then spot the radio signals from these animals so so the answer is yes kind of and increasingly this is becoming more likely so Dirk used the solution in the top right charts you can see some software-defined radio signals so software-defined radio is where you program a hardware device with adaptive frequencies to track the radio pings and then use Python processing so that's the size signals module and others to understand this data the drone takes off in the middle of the jungle and it flies some kind of fixed search pattern and there's no radio contact it comes back to base and the hands then you take the data off of it and you figure out based on the the the radio emissions that it's picked up whether you found any of your signals it's post-processing and you have no control over this this is not a US military drone flown remotely by a pilot this is a fully autonomous vehicle where if it disappears it disappears you have no control over it so it comes back and it tells you where the search part search area should be and I'll give you just an idea of what's if if if this is gonna work let's just see I don't know if you'll pick up the sound that might be on my laptop well if you're missing the sound imagine it going because all it is is a drone taking off so this is a drone taking off you can see Dirk just at the building at the top there looking up he's not controlling it and now he's walking away because this thing's on it's an autonomous flight path I've got a camera hanging down and in a moment it will turn around a bit and then fly off and this is one of the test flights so it's programmed to fly autonomously it's got it's quite a big unit it's about this big fairly deep and it's only got enough batteries and sensor power to get about sort of 30-40 meters of coverage so it flies around a little bit and then it goes off and then this one here it's a four-minute test flight so it flies into the jungle and then it flies back and then it lands and just to prove that it can land so but then it's come back and then it's going back down again and it land so you get the idea that this thing runs autonomously so the first time doke does this they send off one of their units they've done loads of testing he's been testing it in the UK they send off one of these units it flies off amazing great and they're waiting in the jungle they're waiting when they're waiting and they're waiting they're waiting and they're waiting and then it should have come back and they're waiting and they're waiting and then this thing doesn't come back and then they sit back and go right that was the first test flight we have no remote control we have no telemetry the thing has gone and they have to go home and that's it so he flies back to the UK it's pretty disappointing your amazingly disappointing right big drone they've lost in it turns out and when you set the flight path that flies back and forth that assumes there are no obstructions and what hadn't realised was there was a point of elevation where there was a small hill and the trees were a bit higher and you want to fly as low as you can to maximise your signal quality that you'll get to get the widest coverage and they hit a tree and then somebody found this unit one of the the local people found the unit photographed it and sent it back broken very unfortunate a year later they go back into another test and then the second test this was six months ago succeeded so they actually ran their tests back and forth and that plot in the bottom right hand corner is the track of another animal after post-processing that this unit had picked up so now they can more efficiently allocate where humans need to go or the beginnings of the project allocating where the humans need to go to more intensively and accurately track where these orangutangs are to improve their post relocation care and survival and assuming this works you could imagine scaling up the operation saving more animals putting more money into this kind of monitoring project so this was a really nice story of doing good with you data so automate a manual process I would not suggest starting with drones drones are really expensive complicated they break Dirk has horrible stories about trying to get rechargeable batteries into Africa through customs and then you turn up and you realize you the wrong antenna module and everything's wasted and you have to go back and get the right module and then the leg breaks off of the thing and the unit's broken the game so don't start with drones that's really complicated but start with something simpler automate some kind of process and collect your data visualize it and make your decisions so here's the final short story and this hackathon I was involved in one month ago around improving political engagement and you know in the UK we have our ridiculous brexit situation where half the country thinks we need to leave the EU and half the country including many things need to stay in the EU and there's lots of miscommunication there's lots of foolishness it's a real mess right now one of the big problems is most of the populace has become disengaged from the political process and as a consequence and if I choose not to vote then somebody else chooses what will happen for me and that's no good we want more people to be choosing to vote themselves and I realized that it's your election day as well so for those of you who have not yet chosen to vote remember that you do not want somebody else choosing what your future will be like you want the to use what your future would be like so in lunch time go and vote exercise your democratic vote so this is me amplifying the effect of my hackathon from a month ago so the question we asked and I was just one of the attendees at this event the question the organizers were asking was can we get more people registered to vote for the EU election so in the UK it's about 34% of the eligible population will choose to vote so only one-third in our general elections for politics it's two thirds but for the EU elections only one-third so can we get more people register to vote can we using open data only so aggregated mass data not individual data find people who are more likely to have Pro EU sympathies who are more likely to be not registered to vote and then advertise at them on social media to say hey you really should vote with some kind of targeted messaging the group organizing this had loads of different campaigns they had whatsapp systems at surveys and Facebook integration sending good stories around telling people why it's worth getting involved in in the political process my job wasn't to worry about that my job was to lead a team of data scientists at the hackathon to figure out where we might put up which kind of areas or the country put up these politically targeted adverts to say you should vote not who you should vote for but simply you should vote so we got open data from the British government so the second chart shows the turnout in the 2015 general election to one of our local one of our government elections in the UK and we see our turnout there's a percentage there's a chart at the bottom it says sixty and seventy so between 55 and 75 percent of people in different regions of the country were turning out to vote so some of those regions had very low attendance at voting so lots of people not bothering to vote maybe we can engage those people and the third chart shows for the same regions the percentage of women aged 30 to 45 according to the survey some elsewhere in this group women aged 30 to 45 were more likely to have a pro-eu inclination so more interested in EU integration than perhaps an older age group where there is less interest so if we target areas that have a low turnout in the election and a higher proportion of women in this particular demographic we're likely to get more people to register to vote who could be against the EU and for the EU but probably hopefully fingers crossed more likely to be pro-eu but more importantly getting more people registered to vote that's the important thing getting more people to get out and exercise their vote and then how do we make this complex decision of well we've got many regions on the country and many possible ways of deciding which areas to target well we put it on a scatter chart that was the easiest thing we could do so a square scatter charts a bit elongated there all those dots represents one region of the UK if you're higher up the chart that means there was a greater turnout to the last election if you're lower on the charts then it means just a lower turnout and if you're more to the right then it means there were more women in the target demographic who are more likely to vote throw at you so we then took with an if-then rule a couple of the dots in the bottom right hand corner and then they ran an ad campaign off of that and this is the kind of thing where we can't test what the output is we can't do an a/b test to see how well this would have worked we don't know the elections are this weekend in the UK they happen from Thursday so we don't know what the outcome was but I I and a bunch of colleagues used data to try to improve the political engagement process and we were in a group of about a hundred people who were working to improve political engagement in general via lots of different mechanisms so I feel that I did my little bit there to try to make the world a slightly better place so I'm very happy with that and you can too it's let's do a short demo so this is a demo of a different topic entirely we're going to use a jupiter notebook the data scientists will know this the rest of you hopefully you'll see something a little bit new in here with jupiter notebooks and pandas and Seabourn graphing The Guardian is a big newspaper in the UK it's a very big liberal newspaper they put out lots of stories and they spot things that are going wrong in the world and they generally do a pretty good job but the Guardian were all newspapers make mistakes a news article came out by this that this chap tom forese noted this article The Guardian saying there is a collapse in craft beer manufacturing I mean this is a crisis right we've had a boom in craft beer manufacturing in the UK now there's a collapse because big companies are muscling in and wiping out the small companies eight new breweries opened in the last year according to the Guardian down from 390 the year before so this sounds like a massive collapse in the craft beer brewing scene and this is terrible tom forese works with the open data Institute was used to using open data and he uses some power bi to load in some open government data and says hang on a minute these numbers don't make sense so I don't think this is correct so we're going to do something slightly more involved recreating that so this is a Jupiter notebook I've embedded the image of the the story we'll be dealing with oh and Brixton that's not very far from where I live and there's lots of craft beer in the UK I know there's lots of craft beer I enjoy it so can we download the data from the government website CSV file yes of course we can tidy it up and do a couple of graphs so I download the data and I get here 4.3 million rows of CSV data so that loads in a couple of seconds so 4 million rows not big data yes ma'am copies 2 gigabytes I think but it fits into RAM so it's not big data it's small data if it fits into your RAM it's small data that's all you need to gig it's fine we look at the the rows of data with the head command I've got 55 columns so lots of rich data in here about companies particularly the year they were registered in lightly the dates they were registered so I do a bit of cleaning up these are examples of industry codes counted by popularity so I don't know about the industry codes I want to find the beer related industry codes so I do a bit of Investigation so there's other business activities management consultancy information technology and somewhere in this list a much smaller frequency will be the beer related ones and so I I can pull out the beer ones and it turns out there's this 11050 manufacturer of beer this is the one that Tom used in his analysis so I use a data frame query to pull out this particular code that I'm interested in and it pulls me out 2200 rows and these are examples of the companies that are registered and their incorporation dates so we know the the dates that these companies were registered so we can check to see if there's been a collapse recently so if we show a count of in corporations from 2000 to 2018 bigger then we get this and so 2000 bottom left-hand corner there's a small number but a positive number of new beer manufacturing corporations registered and then by 2011 we see a big jump up and then by 2018 the numbers huge so it's just over 300 per year and that's for two years running and then in 2019 we see this strange drop-off why do we see a small count in 2019 at the moment but good is 2019 yeah right so and I'm sure the Guardian didn't didn't do that that would be ridiculous and the Guardian and one of the problems of the Guardian article is they don't quote their data sources know why or how they came up with their conclusions that unquote their data here we've got open data we can point it and say look this is this is pretty obvious right at least it looks pretty obvious to us so we can estimate if it's only been four months of the year and we've got a certain count we can estimate that times by three this will be the estimated count for 29 teams registrations it looks the same as 2017 2018 so I'm not seeing evidence of a massive drop-off that doesn't mean that there hasn't been a significant overnight change in behavior of course and this is where Tom's analysis finished and I said I can thought yeah but what would be another bit of evidence behind all of this and then of course he will ask me her bein after the 2008 global economic crisis wasn't it the case that in the UK we saw a dramatic change in the registration of new businesses wasn't it the case that after many people were put out of work they started their own new companies and indeed that's true so what if we're seeing just an outcome of that effect so if we look at all of the registrations of company it's not just the beer related ones then we see this up and to the right massive magic growth curve this is what startup founders desire to have in their businesses and but this is the the growth of new companies in the UK from the 1950s and in the bottom right hand corner down around here we've got 2000 to 2010 that's when the global global economic crisis occurred and we see some wobbly behavior in new business formation and then just after it shoots up into the riots this is crazy this loads of new business is being registered and this is known in the UK the Bank of England has published reports saying yes after the global economic crisis many people came out of large organisations and started snooze smaller organisations and there were many failures as well of course many of these churn and die very quickly but there are lots of new businesses being formed a new business and your business dynamic occurring in the UK so what if that growth chart we saw a moment ago was simply a reflection of the growth in new business registrations so a way to test that is to get a ratio of every year how many companies are registered and how many beer related companies are registered if the beer related ones are more every year relative to the grossing companies we would see a positive ratio growing over the years but if every year the same number of beer companies were registered as a proportion to the number of new companies we just see a flatline so what do we see we see a positive growth line so that blue jagged line is the ratio of new beer companies to all companies and it goes up and to the right the green line the wiggly one is a smooth moving average that I used in pandas and then that straight line just shows that yes there's a positive inclination from the 1980s but particularly around here so this is after 2010 and it jumps up again so this is after the global economic crisis so I would argue that what we're seeing here is a sustained trend going on over 40 years for new beer manufacturing companies to be increasingly registered every year relative to new business formation in the country so it might be the case that in the last six months we've seen a massive turnaround a significant drop-off and everything has finished but it seems unlikely if you've got a 40 year trend behind something that it turns around overnight so I think this adds extra evidence to the argument that perhaps the Guardian has reported the wrong thing here and that was easily developed using some open data and this is something you could do too so let's finish up the main story I hope you've got that from this is that you should tell useful data stories you have a power for good you have the power of programming most people in the country don't have that you can get some data and do something with it you can make yourself famous within your company and enhance your career by doing this and you can do it in the wider world and either way you'll be making things better so tell useful data stories when I was talking to Ida's last night and I was talking about the the number of pie dates and meetups we have around the world he said well there isn't one here and so it's a I thought it was worth asking the question where's PI data Vilnius would anyone be interested in starting one of those and maybe some of you would and so as Roberta's announced earlier a lunchtime we'll do a data consultation I would call it at lunchtime some of us who do data science will sit down on table hopefully we'll mark it and for any of you who have got questions around your data and what you want do with it come join us and talk and if you can't talk to us talk to your neighbour and see what you can do with your data which tools you might need to answer the kind of question you've got to talk about the questions that you've got that you haven't thought about how you want to answer yet and so if you're interested in talking about that we'll just do that at lunchtime we'll get a table reserved you