Showing posts with label Opendium. Show all posts
Showing posts with label Opendium. Show all posts

Tuesday, 9 March 2021

Making children as safe as they are offline

In a speech last week, The Information Commissioner, Elizabeth Denham said:

The internet was not designed for children, but we know the benefits of children going online. We have protections and rules for kids in the offline world – but they haven’t been translated to the online world.

— Elizabeth Denham, Information Commissioner

Neil Brown from decoded.legal posted an insightful blogpost on the perception that the internet is unregulated and dangerous, compared to the offline world.  The main thrust of the blogpost is that the offline world is not designed to be safe for unsupervised children.

The ICO's Children's Code is intended to make the internet safer for children.  This is a laudable goal, and there are certainly some parts of the code that all companies should be following to protect everyone, children and adults alike.  For example, privacy information is often quite opaque, even to adults, so the requirement to provide clear privacy information would benefit us all.

The offline world is very rarely designed for unsupervised children.  As Neil points out, even children's play areas are usually only designed to be safe for children who are under supervision.  They only prevent unsupervised children from entering by posting a sign (with complex grammar that might not be understood by children).  Washing your hands of the safety of unsupervised children by posting a similar sign on your website or app would almost certainly not be allowed under the Children's Code.

The intent of the Children's Code appears to be to make the internet safe for unsupervised children, but we don't do this in the offline world because it is usually not proportionate.

And this is the crux of the matter: its impossible to make the whole offline world safe for unsupervised children.  It would require banning essential tools, or placing huge financial burdens on vendors.  Councils would have to spend a disproportionate amount of money ensuring that unsupervised children cannot access dangerous roads.  So why do we expect to do so for the online world?

The key point is supervision: the internet should be safe for children, but we shouldn't be going to a disproportionate amount of effort to make it safe for unsupervised children.

But supervision is hard.  If your child is in the kitchen juggling knives, you'll probably notice, whereas they could be in the same room as you, doing unsafe things on their phone and you'll never notice.

Neil briefly points at a few technologies that can be used for child protection:

I use some of the measures which have come in for criticism recently — VPNs, and DNS over https — to maximise the scope of the filtering of Internet connections. More filtering, and more aggressive filtering, not less.

Indeed, I suspect that it is easier to prevent an unsupervised child from travelling to a particular place online than it is offline, if that's the path down which the responsible adult wishes to go.

— Neil Brown, decoded.legal

As Neil points out, by using a VPN you can direct your child's network traffic through system which can block access to inappropriate content and allow parents to supervise their child's online activities.  The parents can install an inspection certificate on your child's device which says "your parents' filter is allowed to decrypt and supervise this device", without allowing unauthorised decryption by others.

This means that the parents, or the child's school, can have a centralised system that they use to set parental controls and supervise the children under their care, across the whole internet.

Unfortunately, the main corporations that control online platforms have unilaterally decided that parents and schools shouldn't be allowed to supervise their children.  In 2016, Google effectively pulled the plug on inspection certificates by disabling them in all Android apps.  Facebook, Twitter and others had already disabled inspection certificates in their own apps some years before.

Without inspection certificates, fine grained filtering and supervision are off the table.  The playground has an opaque fence - you saw your child enter the playground, but you're not allowed to supervise their play.  You know the playground has a slide that is too high for a child of their age, but you're only allowed to control whether they can go into the playground, not whether or not they can go on the high slide.  Are they being bullied in the playground?  Who knows - you're not allowed to look!

There are a few technologies on the horizon which could make it even harder for parents to supervise and control their children's internet access, and historically, Google, Facebook, Twitter, et-al have imposed new privacy technologies and policies upon the public without consultation.  The problem is not the technologies themselves, but that they are unilaterally imposed on users rather than giving them the choice.  Whilst the ability to improve your own privacy is great, the decision over whether a parent can supervise their child should be made by the parents and children themselves, not by untouchable corporations.

The Children's Code does talk about parental controls and monitoring, but there is no framework or requirement to standardise them so that they can interact with a parent's centralised system.  The Children's Code's requirements will simply produce a fragmented approach.  Rather than the parent being able to set controls and supervise their child across the whole internet, they will need to log into each website and app separately.  Imagine having to log in and check separate "is my child juggling knives", "is my child playing with matches" and "is my child bullying their sibling" apps in the offline world.

Rather than demanding that all websites and apps are safe for unsupervised children, the ICO should be setting out a framework for websites and apps to interoperate with centralised systems operated by parents and schools.  They should be placing requirements on companies to consider whether their policies or technologies are detrimental to filters and supervision systems that are already in place.


Note: I am the Technical Director of Opendium, a company that specialises in network based online safety systems for UK schools.  This subject is of importance not only to parents, but to anyone or any organisation that is in a position of loco-parentis, such as schools, foster parents, etc.


Update: 12th March 2021

Neil has posted a follow-up response to this blogpost.

Firstly I'd like to say that, although Neil and I fundamentally disagree on a lot of things, it's very healthy to be having the conversation, and it underscores the fact that there is no single "one size fits all" when it comes to safeguarding children.  Everyone in a position of responsibility over children will have a different opinion on how best to protect those children, and these are the people who should be making the decisions - not governments or corporations, but parents and carers.

Also, although I certainly see technology as a very important part of online safety, I'd never advocate it as the only, or either primary, solution.  Neil is absolutely right that surveillance and supervision are not the same thing, and supervision requires carers to engage with the children and actually teach them how to be safe and to support them.  Indeed, gone are the days where schools just ticked their "online safety" box by installing a filter and letting it quietly run in the corner.  These days, schools are expected to support and teach children to be safe online.  Of course, some schools are very good whilst a few do just install a filter and treat it as a done job.  Thankfully, the inspectors are getting better at asking schools about their online safety policies.  I certainly think that there should be limits on how much carers invade children's privacy, but I also think that children can't expect absolute privacy - there's some balance to be had, and that balance isn't going to be the same for every situation.

My previous comments weren't intended as a rebuttal against Neil's original post - I saw them more as a reflection on something that I think(?) we agreed on (you should supervise children instead of trying to make the world safe for unsupervised children), but our idea of supervision obviously diverges somewhat.  I think this update is probably a rebuttal of Neil's follow up post though.

Traffic decryption

So, without further ado (quotes are from Neil's blogpost):

In most implementations, your target will never know that they are not talking directly to Facebook.

This isn't really true.  Android, for example, has a persistent notification that pops up every time you boot your device reminding you that you have authorised a third party to monitor your connection.  Its not quite as obvious in on a desktop machine, but it is certainly discoverable - clicking the padlock in Firefox clearly shows a warning.  Chrome isn't quite as good, but the information is there.  The persistent Android notification could probably be made more specific, such as telling you who you authorised to monitor your connection, rather than just that someone has been authorised.

State actors have the resources to install certificates directly in the OS's root certificate store, so there's not a lot that OS vendors can do to warn the user about that - this discussion is basically about certificates that the user has authorised themselves.

If someone has built the infrastructure to intercept and inspect your communications in this way, they can look your communications with your bank, the content of your email (and modify it!) and so on.

Entirely true, but thankfully most 5 year olds don't have bank accounts.  I think it goes without saying that how you supervise children depends on a lot of factors.  A primary factor is, of course, the child's age, and what is appropriate for a 5 year old is not appropriate for a 15 year old and certainly not appropriate for adults.  There isn't a "one size fits all" solution, so why should corporations impose one?

Walled gardens

Neil talks about using DNS whitelisting to set up a walled garden that only allows access to specific websites.  This means you to decide which websites to allow access to based entirely on their host name - the rest of the web address is encrypted.  Whilst a great idea in theory, and certainly a staple of school filtering 15 years ago, in the modern age this seems quite naive and doesn't really reflect the reality of the situation for a couple of reasons:

  1. Modern websites use resources from all over the place.  As a recent example, the government's COVID testing website uses Google's reCAPTCHA, which is hosted on www.google.com, so if you wanted to allow access to the COVID testing website, you would also need to allow access to Google web search, Google Images, Google News, Google Videos, etc.  The same is true for most websites and online services these days.  Not only does this undermine the protection of your "walled garden", but it also makes it extremely hard to actually set up the whitelist in the first place - you can't just whitelist the host name of one website, you have to figure out what other resources it needs (this usually can't be automated reliably).
  2. Harmful content is quite often stored along side safe content on the same host name.  If you're allowing access to googleusercontent.com so that various Google applications work, you're also allowing access to a lot of inappropriate content.  Since the child is probably under supervision, it may not be a big concern, but we certainly shouldn't pretend that the problem doesn't exist.

That schools are discouraged from using overly restrictive blocking policies should be an indication that a walled garden approach might do more harm than good.  Parents certainly need to make a decision as to whether its better for children to be in a very restrictive walled garden, or to be allowed to explore the internet more freely with a more dynamic system offering some protection from harmful content they might stumble across.  Again, this is a decision for the parents and carers, not for government or corporations.

Their platform, their rules

The second notion I found particularly interesting was that the private space on these companies' platforms (fixing the weaknesses in their own apps), and the operating systems they develop, should not be theirs to control, and that the decisions as to how they develop their services and products should not be theirs

Businesses, of course, have an obligation to fix security weaknesses in their own apps or platforms (although will I dispute the idea that the user making the choice to allow their communications to be decrypted by a specific party is a security weakness in the operating system).  However, where there are large sections of the population who will be negatively affected by the change, I do believe that a business has an obligation to enter into a discussion to see whether everyone can be accommodated.

Neil's opinion largely seems to be "their platform, their rules, if you don't like it go elsewhere".  But where else can users go?  There are basically 2 choices for mobile operating system:

Android phones start at about £45.  They have the aforementioned problems.

Pretty much the entire online safety sector has been asking Google for a dialogue for the last 5 years and have been roundly ignored.  I've seen numerous bug reports in the Android bug tracker, opened by online safety vendors and schools, and they have all been ignored or closed by the Android team without discussion.

In 2017, the IWF put me in contact with Katie O'Donovan, Google UK's head of Public Policy to try and open a dialogue, but Google were simply not interested in discussing the matter.  People within the Home Office have expressed similar frustrations.

So lets "go elsewhere": an old model iPhone starts from about £300 (£1000 for something more up to date).  Not everyone can afford to spend that kind of money on a phone.

As well as the cost of iPhones, I have to point at an incident that happened around 2 years ago: Apple provides a mobile device management (MDM) system, which is designed to allow businesses to manage their devices.  Parental control software was also allowed to hook into the MDM system, But then Apple changed the rules so that MDM could no long be used for parental control.  Since there was no other system that parental control software could use, there was outcry from the software vendors.  Apple ignored the vendors' concerns and banned the parental control apps from the App Store.  Only later did they reverse this decision as a result of bad press.

So there are only two mobile platforms, and they both have a history of refusing to engage with the people their decisions affect.

If what Steve means is that it should have been left to responsible adults to decide whether or not they want encryption which is MitM'able or not, they do, of course, have that choice: they are not required to let their children send traffic to Facebook or Twitter, if they don't agree with the way they operate, nor are they required to adopt the Android operating system.

The "their platform, their rules" argument could be applied anywhere: Should Facebook be absolved of any child protection obligations, because it's their platform?  Should an outdoor activity centre be absolved of health and safety obligations because it's a private location?  No, of course not - we expect private businesses, both on and offline, to adhere to various duty of care obligations.  Why should we not expect Google, Apple, Facebook, Twitter, Microsoft, etc. to have a duty of care to their users, and to undertake a proper consultation to make sure that changes they make do not undermine their users' safety?  Especially if people have been trying to make them aware of the problems for years.

The idea that we should leave private companies to do whatever they want because parents have a choice to ban their children from those platforms is ridiculous, and at odds with the government's stance with respect to Online Harms.

Surveillance companies, unilateral decisions, and consultations

Lastly, I wonder if there is a degree of double-standards at play here, in that I cannot help but wonder if the vendors of child surveillance systems operate with this degree of transparency and co-operation.

Can these vendors show that those most affected by their software — the children they surveil — were consulted?

No, almost certainly not, but I don't think this is the smoking gun of double standards that Neil wants it to be.  Certainly, as far as Opendium goes, we do not "surveil" children - we merely provide the tools for schools to safeguard the children who are under their care.  Can a CCTV camera vendor show that their customers have complied with the various laws that surround installation and operation of CCTV cameras?  Almost certainly not - in both cases, the vendor is not the company responsible for doing these things, so there is no way for them to guarantee that they have been done.

What I can say is that we do work closely with our customers, and would always advise that they must not undertake any covert monitoring.  Data protection legislation does require schools to be transparent with the children about what monitoring is being done, etc.  I'm not sure why the Information Commissioner's Office has limited the Children's Code to only online services, since much of it is equally relevant to the offline world - schools certainly should be providing clear and understandable privacy information to children.

Do children have a consequence-free option of not being subjected to these surveillance measures?

In law, the child's parents (or the people in loco-parentis) are responsible for making decisions regarding the child's safety.  Do children have a consequence-free option of not being subject to their parent's gaze while playing in the park?  Are they allowed to play in the playground without a teacher watching them?  Probably not - this is not the child's decision, because they are... a child.  It is up to their parents.

But it is certainly a discussion that a child can have with their carer.  I'm certainly aware of one case where a parent requested that their child not be monitored, and the school complied with the request (after having the parent sign a suitable waiver).  I have no idea what the legality of that situation is, given that the result might be the school failing to comply with their statutory obligations.

Anyway, that's enough for today.  As I said at the start of the update, I think these discussions are healthy and, as with politics, we're far better off having a chat about these things to try and understand the opposing point of view rather than just stand at the sidelines shouting "you're wrong". :)

Friday, 28 August 2020

Adventures in Netfilter Land

We're doing some modernisation work at the moment, and part of that is moving our products to a CentOS 8 operating system. Our Opendium UTM appliance includes a firewall with a friendly web based user interface, and the back end is built upon iptables. The end user never sees the iptables bit of course - they just set up some firewall policies based on their users, groups and rule bundles:

Deep packet inspection is done in user space, but for performance reasons most of the other decision making is implemented as iptables rules. Those rules can get pretty complicated, amounting to thousands of iptables rules.

CentOS 8 has moved away from iptables, switching instead to nftables, which is claimed to improve performance, amongst other things. This shouldn't be a big deal - there's an adaption layer to allow nftables rules to be manipulated in exactly the same way as iptables rules, so in theory no need for big changes to our software.

We have plans to move more of the decision making into user space to reduce the complexity of the in-kernel rules, but we don't want to do that right now, so being able to port over the existing system with little modification is great... Except it didn't work.

iptables configurations consist of "chains", where each chain contains a list of rules. A rule is just some criteria, and an action that will be carried out if those criteria match the network traffic. Actions can be things like "ACCEPT" and "DROP" (which allow or disallow the network traffic respectively), or can be a "go to" or "jump" action that points at another chain.

So we can view iptables configurations as a directed graph, with "chains" as vertices and "go to" and "jump" rules as edges. Cycles are not allowed.

Problem 1: "Too many links"

When I tried to load the iptables rules into the Linux kernel, they were rejected with the error "Too many links". Its not a particularly helpful error, but some poking around revealed that nftables has a fixed size stack of 16, which means that your rules can only jump between chains to a maximum depth of 16 jumps.

We do have some pretty complicated rules, but they shouldn't go more than 16 jumps deep, so what's going on?

Delving into the Kernel, we find this horribleness in nft_immediate.c and similar code in nft_lookup.c:

static int nft_immediate_validate(const struct nft_ctx *ctx,
                                  const struct nft_expr *expr,
                                  const struct nft_data **d)
{
        const struct nft_immediate_expr *priv = nft_expr_priv(expr);
        struct nft_ctx *pctx = (struct nft_ctx *)ctx;
        const struct nft_data *data;
        int err;

        if (priv->dreg != NFT_REG_VERDICT)
                return 0;

        data = &priv->data;

        switch (data->verdict.code) {
        case NFT_JUMP:
        case NFT_GOTO:
                pctx->level++;
                err = nft_chain_validate(ctx, data->verdict.chain);
                if (err < 0)
                        return err;
                pctx->level--;
                break;
        default:
                break;
        }

        return 0;
}

This is part of a validation routine that happens when any new rules are added, and is responsible for the "Too many links" error if you try to add rules with a jump depth that would exhaust the 16 frame stack.

We can see that jumps and gotos are both handled the same way - pctx->level is incremented when following either a jump or a goto.  nft_chain_validate() will return an EMLINK error if the level is 16.  This seems wrong - jumps take up stack space, but gotos don't.  Looking at the rest of the code, I can't see a reason for this, so I changed it to the following (in both files), rebuilt, and that seems to solve that problem:

switch (data->verdict.code) {
case NFT_JUMP:
        pctx->level++;
        err = nft_chain_validate(ctx, data->verdict.chain);
        if (err < 0)
                return err;
        pctx->level--;
break;
case NFT_GOTO:
        err = nft_chain_validate(ctx, data->verdict.chain);
        if (err < 0)
                return err;
        break;
default:
        break;
}

There are a couple of other obvious problems with this code which I haven't tried to fix:

  1. ctx is a constant, but the const qualifier is immediately cast away.  Why do this?  Because the author likes watching the world burn?
  2. The "level" attribute of ctx (which is supposed to be a constant) is modified.  It does get restored before the function exits, but that is neglected if there's an error.  I'm guessing that the calling functions just discard the contents of ctx if there is an error, so this probably doesn't really matter, but yuckity yuck!

Problem 2: CPU Lockup

With the above fix in place I tried again and the machine crashed - it totally vanished off the network, and the console was unresponsive and eventually started showing "kernel:watchdog: BUG: soft lockup - CPU#1 stuck for 23s!" warnings.  Great.

When a new rule set is committed, two validation routines are run by the kernel.  In nf_tables_api.c we find the nf_tables_check_loops() function, which checks the graph for cycles and rejects it if it has any; and nft_table_validate(), which calls the nft_chain_validate() stuff, mentioned above, to reject anything that would exceed the stack depth.

These functions are executed each time a change is committed.  Thankfully this is only once per commit, not once per rule - if you use the iptables-restore command, you can make multiple changes at once and just have the lot validated in one go; if you're adding rules one at a time with the iptables command then you only get to make one change at a time so the validation will be re-run for every rule you add.

Unfortunately the algorithm used by these functions is just a brute-force walk of the entire graph, potentially visiting each vertex multiple times.  The following set of rules is a pathological case (it doesn't actually do anything useful, its just a test case to demonstrate the problem):

Chain INPUT (policy ACCEPT)
target     prot opt source               destination         
A0         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain FORWARD (policy ACCEPT)
target     prot opt source               destination         

Chain OUTPUT (policy ACCEPT)
target     prot opt source               destination         

Chain A0 (1 references)
target     prot opt source               destination         
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A1         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A1 (10 references)
target     prot opt source               destination         
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A2         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A2 (10 references)
target     prot opt source               destination         
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A3         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A3 (10 references)
target     prot opt source               destination         
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A4         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A4 (10 references)
target     prot opt source               destination         
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A5         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A5 (10 references)
target     prot opt source               destination         
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A6         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A6 (10 references)
target     prot opt source               destination         
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A7         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A7 (10 references)
target     prot opt source               destination         
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]
A8         all  --  0.0.0.0/0            0.0.0.0/0           [goto]

Chain A8 (10 references)
target     prot opt source               destination         

This is a mere 81 rules, and it takes about 5 seconds to commit them.  Adding an a single additional rule using "iptables -A INPUT -g A0" takes about another 5 seconds because it retriggers the validation, and worst of all, the CPU is locked up for the duration of the validation, leaving the machine unresponsive.

The problem gets exponentially worse, so if we add another chain and another 10 rules the machine will become unresponsive for 50 seconds.  Add another and your server will be AWOL for 8 minutes!

nft_table_validate() always walks the whole graph, nft_table_check_loops() often only walks a subsection of it, but nft_table_check_loops() is actually pretty quick even when walking over the whole graph.  nft_table_validate() is very very slow by comparison.  I couldn't immediately see why there was such a big difference in speed - they both seem to be very similar and neither really does a lot other than recursively walking across the whole graph.

What's the Solution?

The legacy iptables kernel modules still exist in CentOS 8, but the legacy user space tools aren't packaged any more.  But the legacy tools are still part of the iptables SRPM - they are even built, but then omitted from the final RPM.  It was a trivial job for me to edit the spec file, rebuild the package and continue using the legacy iptables modules rather than switching to nftables.

So in the short term, that's what we're doing.  We've got a lot of other modernisation work going on (a big overhaul of the UI and heritable policy structure) so would rather not get side tracked right now.

I've got a longer term plan to switch to using nftables natively, as well as move a lot of the decision making into user space to reduce the complexity of the rules.  We may well end up with a simple enough nftables configuration for these problems to largely go away.  Although I dislike the idea of living with a bug that could basically bring down a server if you load a particularly pathological configuration.

As mentioned above, it's not clear to me why there's a big difference in speed between nft_table_check_loops() and nft_table_validate().  There are certainly better algorithms for the validation they are doing, and even just adding some simple caching to avoid revisiting the same vertices over and over would be a big help.  I may end up revisiting this and rewriting the offending bits of the Linux kernel at some point in the future.

Bugzilla

These bugs have been submitted to the Netfilter Bugzilla at the following links:

Tuesday, 18 September 2018

Chasing wild geese

So, a quick one this morning.  One of our customers has been having problems accessing Lloyds Bank's corporate payments gateway.

The first thing they did was phone Lloyds (very sensible).  Lloyds told them that there were no problems at their end and to clear cookies, add the site to the ActiveX trusted sites list (seriously, why is anyone using ActiveX these days?), etc.  Still not working, so must be a problem with the customer's firewall.

So the customer phoned us.  We pointed a browser at https://payments.corporate.lloydsbank.com/ (on an independent internet connection) and nothing happens - just sits there waiting.  So clearly Lloyds are having problems.

Lets try some slightly lower level debugging:
# openssl s_client -connect payments.corporate.lloydsbank.com:443 -servername payments.corporate.lloydsbank.com
And it just sits there...  Eventually:
CONNECTED(00000003)
And more sitting there doing nothing...  Then eventually:
write:errno=104
---
no peer certificate available
---
No client certificate CA names sent
---
SSL handshake has read 0 bytes and written 0 bytes
---
New, (NONE), Cipher is (NONE)
Secure Renegotiation IS NOT supported
Compression: NONE
Expansion: NONE
---
Oh dear...  I should've got a certificate back from that but instead the web server dropped the connection.  So a few of obvious problems to look into here:
  1. It took a really long time for "CONNECTED" to appear...
  2. It took a really long time for anything to happen once we're connected...
  3. Finally, the web server failed to send us a certificate and negotiate an encrypted TLS session.
Firing up tcpdump, I found that:
  1. There was a significant delay before the first packet (SYN) even appeared.  So something else was going on before the connection was even attempted.  A DNS problem was a good bet.
  2. The first packet (SYN) was resent about 3 times before the web server responded.  This would also cause a significant delay in starting the connection.
So, investigating a potential DNS problem:
dig payments.corporate.lloydsbank.com
Resulted in a long wait before failing - definitely a DNS problem then.  Lets find the name servers responsible:
# dig payments.corporate.lloydsbank.com ns +trace
<snip>

payments.corporate.lloydsbank.com. IN NS  ns-lv6.lloydsbanking.com.
payments.corporate.lloydsbank.com. IN NS  ns-lv2.lloydsbanking.com.
payments.corporate.lloydsbank.com. IN NS  ns-lv7.lloydsbanking.com.
payments.corporate.lloydsbank.com. IN NS  ns-lv3.lloydsbanking.com.
Looking those up gives me:
ns-lv6.lloydsbanking.com. IN    A    141.92.88.1
ns-lv2.lloydsbanking.com. IN    A    141.92.96.1
ns-lv7.lloydsbanking.com. IN    A    195.171.195.169
ns-lv7.lloydsbanking.com. IN    A    195.171.195.168
ns-lv7.lloydsbanking.com. IN    A    195.171.195.167
ns-lv7.lloydsbanking.com. IN    A    195.171.195.166
ns-lv7.lloydsbanking.com. IN    A    195.171.195.165
ns-lv7.lloydsbanking.com. IN    A    195.171.195.164
ns-lv7.lloydsbanking.com. IN    A    195.171.195.163
ns-lv7.lloydsbanking.com. IN    A    195.171.195.170
ns-lv3.lloydsbanking.com. IN    A    141.92.104.1
So there are 11 name servers.  And by trying to look up against each of those name servers I find that 9 of them are down.  That means that about 82% of DNS requests are going to time out - at best things are going to be very slow while the customer's DNS server makes repeated DNS lookups and waits for each to time out; at worst it will fail to find a working DNS server and give up, rendering the website inaccessible.

To summarise:
  • 9 out of 11 of Lloyds' DNS servers were down, resulting in intermittently very slow or even completely broken DNS lookups.
  • If you managed to resolve the web server's IP address, it took a long time to accept the connection.
  • If you managed to get a connection, the web server may fail to negotiate an encrypted TLS session with the client.
With multiple Lloyds Bank servers having serious problems, I wouldn't mind betting that they are being attacked.  But why didn't Lloyds' own support people know / admit that there were problems on their end rather than sending the customer on a wild goose chase - it only took us 15 minutes to diagnose a problem that Lloyds' own people must have already known about.

Tuesday, 31 July 2018

Why do they not listen?

I don't usually talk much about our customers, but sometimes things happen which truly beggar belief.

For many years we have been contracted by a consortium of schools who were geographically close and originally wanted to be able to share a single connection for financial reasons.  This is quite a common arrangement.  Due to the shifting landscape of internet provision, costs and politics, the arrangement came to an end some time ago, which is fine - projects don't last forever, customers come and go and the project ended amicably, with all of the schools involved being pretty happy with us.  In fact, they all ended up taking out independent contracts with us for one thing or another after the project ended anyway.

Just one of the schools has been somewhat... "odd" at times though (I will refer to them simply as "School" to retain their anonymity), and despite our best efforts it has gradually caused problems for them.  Their ICT support is outsourced, which can be both good or bad - some of the companies that provide outsourced ICT services are pretty good, but some of them seem to have a "rip everything out" attitude and insist on unnecessarily replacing a school's equipment on day one instead of spending some time looking to see what is working well and what isn't and only replacing the stuff that isn't working.  This generally seems to be because they want to use systems that they are already used to rather than learning something new.  There's some merit in trying to standardise on systems you know, but obviously leads to a lot of disruption and expense for the school, so in my view is not a great way of doing things.

Anyway, the story probably starts in 2013.

Summer 2013

We had been providing the connectivity between the schools for some years by this point.  Because of limitations of the technology which was available at the time when it was installed, the schools had unreliable, but redundant, interconnects.  These weren't installed by us, but we were contracted to provide and maintain systems to use those unreliable interconnects to provide reliable connectivity.  We were also contracted to provide online safety systems (web filtering, etc.) to the whole consortium.

School: We've just changed our ICT provider and the new provider has decided to replace the existing online safety system with a third party system.
Us: That's fine, but the connections you use for internet access aren't reliable enough to use independently and the equipment you're proposing to remove is used to provide reliable internet access over those unreliable connections. The third party equipment that you want to install is incompatible with the protocols used by the existing equipment, which means you will also need to replace the equipment at the far end of the connections.

Obviously we prefer not to lose a customer, but if they want to switch to another provider then that's fine and we try to guide them and minimise the disruption as best we can.

The ICT provider ordered the third party system, plugged it in, discovered that it didn't work with the interconnects (as we had told them) and ended up backtracking on the whole thing.  So they left our online safety system in place and in fact I think their ICT provider probably decided it was ok in the end - at least they made no further moves to replace it.

Summer 2015

School changed ICT provider again.  This didn't really have any impact on us.

Summer 2016

By now, all of the schools in the consortium, except for School had installed their own internet connections and the interconnects were just being used for a small amount of local traffic and as redundancy in case of failure of one of the internet connections.  The consortium decided that the project had run its course and announced that it would be dismantled by summer 2017.

There was also some discussion about retiring the infrastructure early:

Consortium: We should simplify things by removing all of the equipment that is managing the connections immediately.
Us: We can't do that because that equipment is still need to provide the connection to School.
Consortium: No problem, we'll leave it as-is until summer 2017.

Spring 2017

School: We intend to continue using the existing interconnects after the summer.
Us: It isn't economic to do so since you're now footing the whole bill instead of it being shared by the whole consortium.  The existing equipment is also very old so liable to fail soon.  Replace everything with a new connection, it'll cost less than continuing with the existing equipment.  We can do this for you or you can get a third party in, we don't mind either way.

Summer 2017:

Us: As agreed last year, the existing interconnects will now be shut down.  School will need to migrate over to their new connection.

School: No alternative connection has been procured, there's no time left to get one now, we need to keep using the existing connections.
Us: We already said this was uneconomic, but as a good will gesture we'll take some of the hit ourselves and knock 50% off the cost.  But this is a one year only deal - we will not support this next year because the equipment is well past its end of life.  Also, as the equipment is very old, we recommend you follow our original recommendation and replace the interconnect ASAP since it might fail at any point.  If any of the hardware fails, we won't fix it.

Spring 2018:

Us: Just a reminder, you need to replace the interconnects ASAP.

Summer 2018:

School: We've just changed our ICT provider (again) and the new provider has decided to replace the existing online safety system with a third party system.
Us: That's fine.  As you already know, the connection you are using for your internet access is going out of service this summer, we presume you've procured a replacement?
School: No we haven't, we intend to continue using the existing interconnects.
Us: But those interconnects aren't reliable enough to use independently and the equipment you're proposing to remove is used to provide reliable internet access over those unreliable connections. The third party equipment that you want to install is incompatible with the protocols used by the existing equipment, which means you will also need to replace the equipment at the far end of the connections.
Us: In fact, this is exactly what we said in the summer of 2013, then again in the summer of 2016, then again in spring 2017, and in summer 2017, and in spring 2018.
Us: Also, we already told you a year ago that we weren't going to support any of this equipment which manages those interconnects any more as it is far too old.
School: We don't need that equipment, we're just going to use a single (15 year old, unreliable) connection in isolation.
Us: Errm, that will be really unreliable, here are some statistics from our monitoring data to show just how unreliable it will be.

Panic ensues at School.

School: Why didn't you tell us that we would need a new connection!  It's now far too late to procure one in time.  You must extend our contract for free.
Us: Umm, no.
Us: We've made every effort to recommend the most economic and reliable way forward and have been ignored at every step.
Us: Last year we dug you out of a hole you'd made for yourselves and even gave you a big cost reduction out of our good will.  You have repaid us by continuing to ignore our recommendations, blaming us for the mess you've got yourself in, paying your last invoice months late and cancelling your contract with us.
Us: You have now demanded that we dig you out of a hole again out of the goodness of our hearts at extremely short notice.
Us: Here are a selection of get-out-of-jail cards from our standard price list, which we are happy for you to buy from us at the standard price, go pick one.

And apparently this is all our fault...  At least now that the contract with them has ended we won't have to deal with any of the fallout from this mess.  Seems like a classic case of "I think we've heard enough from the experts" to me :)

Friday, 15 July 2016

Adventures in broken apps

Its been a week of frustration, but also success.  We always have a perennial problem of broken apps, which work ok on the average home network, but as soon as you try to secure things you hit problems.  Although there is a good argument for keeping school networks fairly permissive, realistically schools can't just turn off their firewall and let everyone at it - not only would the school be failing in their duty of care, but it would be a security nightmare too.

We're frequently tasked with getting some badly behaved app to work on a school network.  Unfortunately when something works elsewhere but breaks when connected to the school network, the firewall/web filter is often regarded by staff to be "broken", even when we can clearly see that the app is the one doing broken things.  Still, we like to keep our customers happy and be as helpful as possible.

We routinely spend a lot of time diagnosing problems and sending debugging to the app vendors and there are a few vendors that are thankful for our input and will work with us to improve their software.  However, I think its fair to say that the vast majority of app vendors are completely uninterested in fixing bugs in their software.  This attitude is unfortunately prevalent across all kinds of suppliers - from small suppliers, right up to the likes of Microsoft, Apple and Google.  In fact, we no longer submit bug reports to Apple because collecting data for them uses a huge amount of our engineering time and they have never fixed any of the bugs we've reported.  A good example of this recently was CloudMosa, who responded within 24 hours of our bug report, explaining that they weren't going to fix the bugs that we reported in their Puffin Academy app.  WhatsApp have been similarly unhelpful with problems we reported, stating "WhatsApp is not designed to be used with proxy configurations or restricted networks, and we cannot provide support for these network configurations."  What a cop-out!

So with the app developers refusing to properly support their own software, our customers have nowhere left to turn and it is often down to us to do our best to work around the flaws in the applications.  A lot of this comes down to collecting as much information as possible about each connection so we can automatically turn security features on and off on our systems to work around incompatibilities.

We do things like snooping on TLS handshakes - when a device sets up an encrypted connection, it is supposed to include information such as a "server name indication" and we can spot that and use the presented server name to decide whether or not its safe to intercept and analyse the encrypted data.  Some apps aren't compatible with interception, so when we see known problem apps we avoid intercepting their connections.  Unfortunately, every so often you find an app that doesn't bother to include this data, and there's no way for the system to know if it's ok to intercept the connection.

In the second half of this week we developed some new code for the web filter to actively callout to the remote server in these situations.  When the web filter sees an encrypted connection that has no server name indication, it can now connect to the web server, retrieve the certificate and use information in it to figure out what to do next.  We're expecting this to help a lot with the problem apps.  The results of each callout are cached to reduce the impact on the web servers.  This is currently going through testing to make sure it won't cause any problems,

Another frustration has always been Skype - this has always been a real pain to make it work reliably and securely.  We've spent a lot of time this week pouring over network traffic dumps and testing.  There are numerous problems with the Skype protocol, which boil down to:
  1. It makes connections on TCP port 443 (and therefore look like HTTPS), which aren't actually HTTPS, or even TLS.  These connections can go to any IP address, so we can't trivially poke holes in the firewall for them.  They get picked up by the transparent proxy, which treats them as encrypted HTTPS connections and therefore fails to handle them since they aren't actually HTTPS.
  2. It makes real HTTPS connections carrying WebSockets requests.  Unfortunately we don't yet support WebSockets and as Skype doesn't bother to include a server name indication we can't pre-emptively decide not to intercept them.
  3. It sends peer-to-peer voice and video media over UDP using any port numbers between 1024 and 65535.  Since it's peer-to-peer, this traffic can be directed at any IP address on the internet.  Official advice is to just allow that through your firewall - if you do that you may as well not even bother to have a firewall in the first place!
  4. All of Skype's traffic is encrypted so it's almost impossible to figure out what it's actually trying to do and what went wrong when it fails.
  5. If something goes wrong, Skype just breaks in one way or another and provides no indication what actually went wrong.  The Android version of Skype could output some debugging data to Android's standard debugging log, but it doesn't.  The PC version of Skype can be told to produce a debug log, but the log is encrypted so that only Microsoft developers can read it (gee, thanks for nothing Microsoft!)
Fortunately, it turns out that the not-HTTPS-that-looks-like-HTTPS (1) traffic isn't needed if it can successfully set up the peer-to-peer UDP connections (3), so we think we can ignore that problem.

It doesn't actually seem to matter too much if the WebSockets connections (2) fail, but this should be handled by the web filter's new TLS callouts system described above.

So we're left with the UDP traffic, which can be to an IP address on any port (3).  This one is a real problem - blindly allowing all of this traffic would also allow a whole load of other stuff such as VPNs, games, etc.  So we've been playing with the nDPI deep packet inspection library and nDPI-Netfilter.

Normally, firewalling is done based on just the information in the packet headers, such as the source and destination addresses.  Deep packet inspection examines all of the data associated with the connection, including the payload, in an attempt to identify what protocol is being used.  We seem to have got this working pretty reliably now.  The sticking point is that the deep packet inspection system needs to see a few packets before it can identify the protocol - usually you'd allow or refuse the connection immediately, but for DPI to work you have to allow all connections for a while and then terminate any that you don't want to allow.  We're finding that allowing the first 10 kilobytes seems to work reasonably well - after that we chop any connections that haven't been identified as Skype.

Of course, all this was massively complicated by the fact that, unbeknown to us, Skype had a bug which made video unreliable - we found that out on Wednesday when Microsoft released a new version to address the problem.  But not before spending a lot of time trying to figure out what was going wrong (did I mention that Skype problems are almost impossible to debug because absolutely everything, including the debug log, is encrypted so you can't examine it?)

The original intention was to implement deep packet inspection in the new firewall system which we are developing, but by popular demand we've backported this to the existing firewall.  There is currently no user interface to set up the Skype DPI rules, but we can manually set them up for customers on demand for the time being.

Anyway, a moderately successful week - we're still testing the Skype rules, but they should be available Real Soon Now™.

Friday, 22 January 2016

Performance improvements

I've been doing quite a lot of work on improving the performance of our Iceni web filter.  An increasing number of schools are now getting internet connections exceeding 100Mbps, and there was a noticeable drop-off of performance at higher throughputs.

We had previously identified a number of areas which could have been acting as performance bottlenecks- amongst these were reduction of the number of memory copies by replacing linear buffers with ring buffers (using mmap() tricks to present the rings as linear memory) and replacement of some of the content analysis code.


However, in testing, we found that the CPU didn't seem to be the limiting factor.  Even when going flat-out, there was plenty of spare CPU time, and we just weren't seeing the performance we expected.  This was unexpected - we had thought that performance would be harmed by inefficiencies, but if that were the case, we'd expect to see all of the CPUs pegged.

Eventually this was narrowed down to a locking bug.  The software is multithreaded - this means that there are effectively multiple copies (threads) of the program running at the same time, all accessing the same data in memory.  The data is "locked" while a thread is accessing it so that another thread doesn't come along and change it.

Its safe to have lots of threads reading a piece of data at the same time, but we definitely don't want a thread to change that data while other threads are accessing it.  So threads can either make a "nonexclusive lock" or an "exclusive lock" - multiple threads can hold nonexclusive locks at the same time, but if a thread holds an exclusive lock, no other thread can acquire a lock (either exclusively or nonexclusively).

So when a thread asks for an exclusive lock, two things have to happen:
  1. It has to wait for all of the other locks (exclusive and nonexclusive) to be released.
  2. No new locks must be acquired by other threads while it waits.
In our code, when a thread is waiting for an exclusive lock to be acquired, it sets an "exclusive" flag, and once the lock has been acquired, that flag is cleared.  Threads trying to acquire nonexclusive locks check this flag and the nonexclusive locks are therefore inhibited.

The problem arose when multiple threads were waiting for exclusive locks at the same time - they would all set the "exclusive" flag, but the first one to successfully get the lock would clear it again, even though other threads were still waiting for exclusive locks.  Nonexclusive locks where therefore no longer inhibited, and the threads trying to get exclusive locks would be competing with them.  The result was that it could take an extremely long time (frequently hundreds of milliseconds) to acquire an exclusive lock!

The fix was simple - the "exclusive" flag was replaced with a counter, which is incremented when a thread is waiting for an exclusive lock and decremented again when it acquires the lock.

This fix has been rolled out to a number of customers and the user experience improvement has been striking - not only can significantly higher throughput be achieved, but the responsiveness of websites is noticably much better, even at low throughputs.

Going forward

We've already done a lot of the content analysis code improvements, although there are still some more to come.  The main thing now is replacing the linear buffers with ring buffers, which will reduce the amount of memory copying needed.  This is involving a rip-out and replace job on the ICAP protocol interface and finite state machine.  The new code is looking a lot neater and much easier to understand, and is giving us the opportunity to better optimise memory usage based on what we've learnt in the years since it was originally written.

The new work revolves around a neat ring buffer library I wrote a few months back - using mmap() tricks, the buffer is presented to the rest of the application as linear memory.  Not having to worry about data crossing the start/end of the ring simplifies things a lot, and this library has already been used in anger elsewhere in Iceni, so it can be treated as well tested.

Thursday, 17 September 2015

Ranting about LEA Network Administrators

I'm getting increasingly tired of the network administrators at a certain LEA.  I'm going to venture that they aren't really qualified to run the LEA's WAN...

Back at the start of July, one of our customers reported that an application was intermittently extremely slow or completely failed.  Originally the customer thought that it was a firewalling problem, but we identified a DNS problem as the cause - the LEA's DNS server was taking a few orders of magnitude longer than you'd expect to respond to AAAA record lookups for two domains that were used by the app, and eventually responded with a failure.

A quick explanation is probably in order here: when a client needs to connect to a web server on the internet, it has to convert the domain name (e.g. www.example.com) into the numerical IP address(es) of the server(s).  It does this through the Domain Name System (DNS).  Typically a (modern) client requests both "A" records and "AAAA" records from a DNS server - "A" records list the (legacy) IPv4 addresses for the web server, whilst "AAAA" records list its (newer) IPv6 addresses.  A web server may not have both IPv4 and IPv6 addresses, but the DNS server still has to produce a successful response to tell the client this.

Importantly, even if the client only has a legacy IPv4 internet connection, it may not know that it can't contact a server using IPv6 until it actually tries, so it will usually still ask for "AAAA" records so that it can get an address and try it.  Also, whether or not the DNS server has an IPv6 connection is irrelevant - if AAAA records exist it is required to reply with them, and if they don't its required to reply saying they don't; the LEA DNS server was doing neither.

The LEA were informed that their DNS server was breaking when queried for the AAAA records, so they replied saying they had pinged a few things and used done some nslookups (presumably only for the A record!) and they couldn't see a problem.

We ran more tests - looking up the A records worked fine (successful response in 9 milliseconds), looking up the AAAA records failed (failure response in 17 seconds) and looking up the AAAA records through a different DNS server (successful response in 23 milliseconds).  We even gave them transcripts of the tests so that they would know exactly how we tested it and would be able to reproduce it themselves.

The LEA responded with words to the effect of "well no one else has reported this problem", so we ran the same tests from a different school within the same area and demonstrated that they had the same issues.  Again, we sent transcripts of the tests to the LEA (* see footnote).

The LEA then started asking whether the school was using a transparent proxy and what the school's internal domain name is - none of this is relevant to the problem being reported.  We weren't reporting problems with the transparent proxy, or any of the school's internal servers, we were specifically reporting a problem with the LEA's DNS server.

We did some further investigation and got more detail on which DNS lookups were failing, sent this to the LEA together with more transcripts of tests and an offer to work with them to help.  Rather than asking for our help, the LEA closed the ticket as "resolved", but provided no explanation.  We reran the tests, sent them another transcript demonstrating that nothing had been fixed.

The problem was originally reported at the start of the summer holidays.  Two months later the new term started - still the problems weren't fixed, still the LEA hadn't taken us up on our offers to help them (for free!) and now it transpires a lot more domains are affected than we originally investigated.  Its causing really serious problems for the school, so the school started banging heads together and someone from the LEA actually called us.  I explain the problem yet again and he goes off saying he needs to look up some more information.

Then they start talking about transparent proxying again, and again I have to point out that we are reporting a problem with the DNS server and that this has nothing to do with the transparent proxy.  Again, I send them an email describing the problem, providing transcripts of tests, etc.  LEA techie tells the school that I didn't send any information and that I just forwarded his email back to him - I'm a bit stunned about this since it means that (1) he has never seen an email with inline comments before, and (2) he didn't read past the first line of the email.  So the email gets resent to him.

The LEA reply with some screenshots of some tests they have done which they say show that there's no problem:
  • They logged into the leased line router and queried the network interface statistics that show no line errors.
  • They pinged a few machines.
  • They tracerouted to somewhere.
i.e. they didn't test the thing we actually reported being faulty.

The LEA suggests that this is happening because they don't provide IPv6 connectivity (as mentioned above, whether or not IPv6 is available doesn't actually change anything from a DNS perspective - clients still look up AAAA records and DNS servers are still expected to reply).

Now they say they've poked lots of holes in their firewall because they "have no information on what port AAAA records would be using" (errm, 53, the same as every other DNS request in the world?!) and could we retest - unsurprisingly its still broken.

As far as I can see:
  1. They haven't actually run the tests (which we've told them how to run!) to try and reproduce the problem.  They've tested a few other things that were never a problem to begin with.
  2. They don't understand enough about DNS (which is an extremely fundamental internet protocol) to diagnose the issues - they seem to have entered a "change something at random and see if it fixes it" phase instead of trying to get to the root of the problem.
  3. They are completely out of their depth - if they want to run a reliable WAN, they need someone wuo is actually qualified to administer a network.  That means someone who understands how to reproduce problems, use debugging tools such as WireShark, etc.
  4. They haven't handled this in a timely way at all - they had the whole of summer to investigate, and didn't actually start looking at anything in earnest until after the start of term.

I have spent literally hours on this problem, mostly repeating the same explanations and tests over and over (although strictly speaking this isn't "our problem", diagnosing and liaising with the LEA is something we're handling as part of the customer's advanced support contract, so we're not really being paid by the LEA for this level of hand-holding).  I honestly can't see them resolving this problem until they reproduce it themselves and do some proper diagnostics.


Footnote

As mentioned, part of the LEA's defense is basically "no one else has reported a problem" - now, not looking into a problem because it isn't affecting many people is a pretty crumby attitude to begin with, but there are reasons why some people would be affected and some not.

Fundamentally, how services, such as DNS, are expected to behave are defined by standards.  These boil down to rules like "when a client sends a request like this, the server must send a response like that".  Software that relies on these services is written to expect them to follow the rules laid out by the standards, and there is no standard set of rules saying how to handle a service that is breaking the rules - it is extremely difficult to draw up a standard explaining how to deal with something breaking the standards, simply because there are so many ways the standards could be broken!

So you may have two different pieces of software that do basically the same job, call them A and B.  In an environment where everything is sticking to the rules, they both work equally well since this behaviour is standardised.  However, if some service isn't sticking to the standards then they will often handle this differently - maybe software A still works fine, but software B breaks.  In a different situation the roles may be reversed, with software B working ok.

So its possible for a real problem, such as this, to go unreported simply because a lot of people happen to be using software that, by chance, isn't badly affected by the broken service.  Its also possible for problems to go unreported because people write off the problem as "software A is broken" and so don't report the issue to the operator of the broken server.

Tuesday, 4 August 2015

Counter-Terrorism and Security Act 2015

Firstly, lets get a disclaimer out of the way: I am not a lawyer, none of this constitutes legal advice, etc.  I am also of the opinion that the threat of terrorism is minuscule and that if all the effort the government puts into anti-terrorism were instead put into road safety or health care, a lot more lives would be saved, but anyway...

It has come to my attention that one of our competitors has been engaging in a bit of scare mongering in an effort to sell their product.  The following email from them has been going around a number of schools:
As of July 2015, schools across the UK are subject to a duty under the Counter-Terrorism and Security Act in which they are required to have "due regard to the need to prevent people from being drawn into terrorism". This duty is known as The Prevent Duty.

[We have] been working with various governments around the world for over a decade, developing solutions to help schools and colleges protect their students from potentially harmful sites and information. All of our solutions have been developed purely for education, based on feedback from IT Staff, teachers and IT Professionals to ensure that they have the tools they need to prevent and report on dangerous activity.

Our Web Filter is a fully customisable, cloud based solution that allows granular filtering and reporting with great ease. The Web Filter includes Suspicious Search Engine Queries, Internet Lockouts and real time updates.

Click here to learn more or contact us directly.
The Counter-Terrorism and Security Act 2015 is a pretty long and convoluted bit of legislation, but thankfully there's a Prevent Duty Guidance document, which is significantly easier to read and provides more specific guidance on what institutions are actually expected to do.  The following is a brief analysis of the parts of the guidance that relate to ICT operations within schools - there are numerous non-ICT responsibilities listed in the guidance which I won't cover here.

The guidance immediately makes clear that no "new functions" are conferred upon a school (paragraph 4):  You don't have to do anything you weren't already doing, you're just expected to place an appropriate amount of weight on preventing people from being drawn into terrorism.

There is no specific mention of monitoring internet activity, although there are several suggestions that internet filtering should be considered (paragraphs 45, 71 and 97).

A passing mention to having "effective IT policies in place which ensure that [signs of radicalisation] can be recognised and responded to appropriately" is made (paragraph 79), but there are no specific policies suggested.  Institutions must have clear policies relating to the use of equipment, especially with regards to legitimate research into terrorism/counter-terrorism as part of learning (paragraphs 97-98).

Institutions should develop an action plan to set out any actions that they will take to mitigate the risks (paragraph 90).

In short, most schools already have an existing filtering solution and robust policies regarding the use of equipment, and that is really all that is required.  (And if they don't already have this, they should, for many reasons besides this legislation.)  There certainly seems to be no requirement to replace existing systems, unless they have unusually poor capabilities.

Comparing our Iceni product with the competitor's offering, we think our customers are actually in a better position to protect their students and staff than our competitor's customers.  Not just protecting them from being drawn into terrorism, but from many other risks too; and protection is surely far more important than just meeting the minimum requirements of the legislation.

It seems that only minimal work is required to comply with the legislation - e.g. the "action plan" for risk mitigation should be drawn up and probably include information about what reports the ICT staff should be running on a regular basis, and what they should do if they find anything concerning, etc.

As always, we're very happy to work with customers to resolve any concerns, to help investigate suspicious activity and even compile data for the police.

Wednesday, 15 April 2015

Following scripts

One of the most annoying things is contacting technical support to get a problem solved, and getting answers that have clearly come from a script, rather than being able to talk to someone who can actually put some thought into the problem.

https://secured.studentfinance.direct.gov.uk is using an expired certificate.  Here's the output from OpenSSL:
depth=2 C = SE, O = AddTrust AB, OU = AddTrust External TTP Network, CN = AddTrust External CA Root
depth=1 C = GB, ST = Greater Manchester, L = Salford, O = COMODO CA Limited, CN = COMODO SSL CA
depth=0 OU = Domain Control Validated, OU = COMODO SSL, CN = secured.studentfinance.direct.gov.uk
verify error:num=10:certificate has expired
notAfter=Apr  2 23:59:59 2015 GMT
All pretty straight forward - the certificate clearly expired at the end of April 2nd.  This causes web browsers to display very prominent security warnings (not just some easy to ignore icon somewhere - you actually have to click through the security warning and tell it that you understand that you're taking a risk).  So being the good citizens that we are, we informed them so they could fix their systems - here's the response we got back:
I would advise to clear all browser cache, cookies and temporary internet files, close the browser and reopen a new one.
How exactly does that address the problem we raised?  It doesn't does it - clearly their support department works on the principle of "don't bother looking into the problem, just tell people to clear their cookies and hopefully they'll go away".

I'm glad that we've never gone down that route with our customer support - when people phone us they don't get fobbed off with some scripted reply; they actually get to talk to someone with a lot of technical experience who can start looking into the problem immediately and provide a relevant answer.

Wednesday, 26 November 2014

Are exclusivity deals good for the consumer?

Opendium has been going for over 9 years now, and over that time we've gained a number of schools as very happy customers through word of mouth (and as a testament to their satisfaction, no school has ever left us!)  We've spent that time working with our customers to build a very capable product, which is also somewhat cheaper than most of our competitors, and we're actively working to promote our product to more schools.

We're primarily marketing to independent schools, since they have much more freedom to make their own decisions, and there are a number of organisations that represent the British independent schools which we're actively engaging with.  This should be good for both us and their members.  Just last month we sponsored and attended the Welsh Independent Schools Council's annual conference, which was a good experience for us and brought up some interesting ideas for our product roadmap.  In fact, we've already implemented some of those ideas, and they are now going through the testing phase of our development cycle.

However, I was surprised by the Independent Schools Association's attitude when we contacted them - they refuse to work with us as they say they already have exclusive contracts with "preferred suppliers".  They are supposed to be working for their members, which seems at odds with any kind of supplier exclusivity.  Competition is almost always good for the consumer - it leads to innovation and lower prices, and conversely exclusivity almost always leads to stagnation and high prices.  Surely if they truly are working for the interests of their members, they would be trying to foster as much competition as possible and giving their members a wide variety of suppliers to choose from to meet their individual needs?

Thankfully, ISA's attitude doesn't seem to be shared by the other people that we have been in discussions with, so we're looking forward to working with them and benefiting their members.

Monday, 13 October 2014

Small suppliers picking up the tab for support

I've been having some thoughts about how the support load is distributed between large suppliers such as Microsoft, Apple and Google vs. smaller suppliers such as ourselves.

It seems that people buy in services from big providers, but when problems hit they can't get the necessary level of support from them, so turn to the smaller providers with whom they have a not-entirely-related support contract.

This happens to us all the time - we have all manor of kludges in our software to make Apple devices work reliably with it due to bugs in Apple's software, for example.  In an ideal world, the customer would call Apple and Apple would diagnose the problem (possibly with our help) and fix their software.  In a less than ideal world, we would do a temporary work-around to get our customers up and running, then report these bugs to Apple who would fix them.  Back in the real world, Apple never fix the bugs and the work-arounds become permanent bodges that are an ongoing minefield for us.

We used to report all the Apple bugs we found to Apple with the expectation that they would be interested in fixing them.  These days we don't bother - they have never shown any interest in fixing a bug we've reported.  Usually reporting went like this: after spending hours diagnosing a problem, we send them a comprehensive report of what's happening, how to reproduce it, often with network traffic dumps clearly showing the problem.  They respond asking for exactly the same information as we just provided, but in a different format.  So we spend hours reproducing the problem again, send them everything they asked for and never hear anything back.  I've got no problem with spending some time collecting information for them if they are actually going to use it, but it's a complete waste of our time to do debugging that they will ignore every time.

We're currently having problems with Microsoft's web servers - they have a bug in them that means certain clients can't connect (pretty much anything using OpenSSL on Scientific Linux 6.5).  In particular, our proxy server software can't connect without some work-arounds.  There is no well publicised address for reporting bugs, but we found a promising looking address and sent a comprehensive bug report.  We even prefaced the report with a "if this is the wrong address, please forward it on to the right department" note.  Instead, we have simply been bounced from department to department, many refusing to hand out email addresses and instead insisting on us phoning - the phone operators are completely ill-equipped to handle this kind of bug report and inevitably bounce us on to another department.

So much like Apple, Microsoft seem disinterested in actually looking at bug reports.  Limited experience of dealing with Google is much the same - they are just too big to be interested in resolving problems that don't affect hundreds of thousands of customers.

So we're back to my initial thoughts - customers buy expensive products from the big guys, who pocket the profits and refuse to support them properly.  Leaving us to have to pick up the pieces despite it not really being part of our remit, because telling a customer "we're not going to help you" isn't really an option for us.  Yet somehow, when we are unable to work around the problems it is somehow seen as our fault and reflects badly on us - no one ever stops buying from the big guys because of this stuff.

Thursday, 11 September 2014

Diagnosing Sharepoint Breakage

Every so often you get a proper puzzle to solve, and this morning is one of those times. One of our customers reported that they were unable to contact the Microsoft Sharepoint servers through their proxy server, a quick test on my test system confirmed the same issue so we spent about 3 hours delving right into the nitty gritty to figure out what was going on.

The proxy was reporting "connection reset by peer" during the TLS handshake - TLS (Transport Layer Security) is the cryptography protocol used to secure HTTPS web sites, and TLS problems tend to be a pain since the OpenSSL library usually doesn't give especially verbose error messages.  It was clear this wasn't going to be a trivial problem to solve so we immediately disabled HTTPS interception for the Sharepoint site to get it up and running again.  Customer confirmed that this had resolved the issue, so that takes the pressure off a bit but raises a question: why is it working ok when the browser is negotiating the encryption, but not when the proxy is negotiating?

The first port of call was to capture some network traffic and load it into Wireshark for analysis.  This showed that the proxy is sending a TLS "Client Hello" handshake, the server was returning a TCP ACK, but no TLS response.  30 seconds later the server tears down the connection with a TCP RST.  The ACK confirms that the server got the "Client Hello", and you'd usually expect the response to be sent in the same packet as the ACK so it looked like the packet wasn't being dropped by intermediate network hops - the server simply was never sending a handshake response.

Time to make things simpler - instead of using the proxy server, lets ask OpenSSL to connect directly:
openssl s_client -showcerts -connect 157.55.229.87:443
This failed in the same way when we tried it on the test server, but succeeded when run on my Fedora workstation.  Comparing the network traffic between the working and non-working tests showed that the most obvious different was that the non-working handshake presented a few more ciphers for the server to choose from - maybe one of those extra ciphers was confusing the Sharepoint server.

We tried adjusting the list of cipher suites, but each time we tried we found that the request succeeded and we couldn't pin down anything specific that would break it.  We needed to start with the broken handshake and edit it bit by bit until it started working - that would let us figure out specifically what needed to change to make it work.

So we took the captured network traffic and dumped it out as hex:
tcpdump -r capture.pcap -x > capture.hex
We're not interested in the TCP layer stuff, so the first three packets can be ignored (SYN, SYN ACK, ACK) - these are the normal TCP three-way handshake.  The next packet contains the "Client Hello" which we're interested in, but it also contains the Ethernet, IP and TCP headers.  Using Wireshark it's trivial to identify the start of the payload, and we just trimmed everything before that off the hex dump.

Now to replay it and make sure it still fails:
(sed -e 's/#.*$//' capture.hex | xxd -r -p ; sleep 5) | nc 157.55.229.87 443
The sed bit at the start just strips off anything after a # so we can put comments in the hex file.  xxd converts it back into binary and we used nc to connect to the web server and send the data.

We checked the traffic in Wireshark - all looks as expected and the web server still didn't respond, so far so good.

Again, using Wireshark we can identify the various parts of the packet, and set about modifying them.  Of interest are four headers indicating the length of various sections - the TLS Record Layer has an overall length header, within that there is the "Client Hello" data which has its own length header, and within the "Client Hello" are a cipher suite list and an extension list, which again have their own headers indicating their respective lengths.  Each length header is 16 bits long, so can contain a value of up to 65535.

As mentioned, we were interested in the cipher suites - in particular the extra ones that were presented in the broken handshake but not in the working one.  So we set about removing them one by one - each cipher suite is 16 bits long, so removing it involves deleting it from the cipher suite list, and then reducing the cipher suite length, client hello length and tls record length headers by 2 each.

Each time we removed a cipher suite, we replayed the data to the server and looked to see what happened.  After removing two cipher suites, the server suddenly started responding with a "Server Hello"!  We put these ciphers back and removed two others so see if it was specifically one of those ciphers confusing the server, but that didn't break anything again - the server was still happy.

The broken handshake that we started out with had a TLS record length of 258 octets and removing two ciphers (16 bits each) reduced it to 254 - a number that will fit in a single octet, whereas 258 requires two octets.  So we tried adding all the ciphers back in and removing one of the records from the extensions list (5 octets) instead.  Again, the server responded and was happy.

So there we go.  It looks like Microsoft's Sharepoint server has a bug in it that breaks any client that tries to handshake with a TLS record more than 255 octets long.  Evidently the proxy presents a larger selection of cipher suites to the server than most web browsers, so it works fine from the browser but not from the proxy.

We have contacted Microsoft, although I have no idea if we've contacted the right department but hopefully it will get passed on to the right people.

IP Based Controls

Over the years, our Iceni servers have undergone a number of design changes in order to accommodate the changing nature of devices being used on networks.  In particular, authentication of web traffic has needed special attention constantly.

In the old days, software usually had really good support for authenticating with web proxy servers.  Windows clients would silently authenticate each web request using NTLM or Kerberos, non-Windows stuff used HTTP Basic authentication (it pops up a username/password box when you start a session, but you can use the "remember password" checkbox to stop this getting annoying).  Every so often we came across a rare example of software that couldn't handle proxy authentication and we'd have to tweak the proxy configuration a bit to bypass the authentication and filtering, but in general life was good.

We're increasingly seeing software support for web proxy servers getting poorer though - quite a lot of software just plain ignores the system-wide settings and bypasses the proxy, and an increasing amount of software can't handle proxy authentication at all.  In fact, this latter point has often shown just how poorly built some of the modern software is: Windows 8, for example, tries to log in to login.live.com when you log into a machine, and if the proxy asks for authentication the machine hangs and has to be hard-reset!  Apple devices seem particularly bad too - if the proxy asks an iPhone to authenticate when it tries to synchronise its calendar, it just retries immediately, and keeps going indefinitely - hundreds of times a second!

A few years ago we did a lot of work to work around these broken devices:  Firstly we introduced a transparent proxy to deal with the software that completely ignores the proxy settings.  Then we started caching the last known user for each IP address and for client software that's known to be broken, we just reused those details rather than asking them to authenticate.  We also added a captive portal and support for WISPr authentication to help the Apple devices along a bit.

Along the way we his some surprising problems - for example, you would expect every web request to be independent of each other, but we found that if we avoided authenticating certain iPhone traffic, then completely unrelated traffic from that device that would usually work fine suddenly stopped being able to cope with authentication too!

The move to Iceni 2 saw more changes - administrators can now tell the system to only use the captive portal/WISPr authentication for certain problem URIs, or disable authentication entirely in some cases.  For example, by default Iceni 2 servers don't authenticate of filter login.live.com.

All this work has gone a long way to avoiding the problems that were cropping up, but increasingly there's a feeling that things like phones and tablets have such poor support for HTTP proxy authentication that its probably preferable to turn it off entirely for those devices and rely on the captive portal and WISPr.  But how do you do that just for those devices, and not for things like the on-domain Windows machines which still work fine?

This brings me on the the latest stuff I've just finished working on and is now going through QA testing (soon to be released to the customers, all being well!):  We now allow you to define a network - a network address and netmask - and drop it into a user group as if it were a user.  This means you can do stuff like disabling authentication for all devices on a particular network - your wifi network, for example.

The bonus of this is that, if your network is split up appropriately, you can also tweak filtering based on the workstation's location - you can relax the filtering for class rooms that are well supervised, for example.

We've also got rid of the "Guest" user and renamed the "Guests" group to "Anonymous" to better reflect what it means.

Over all I really like the new model, and I have plans to extend it to the mail server component as well.

However, my new job for today is to fix a locking bug in the web filter - joy of joys!