Our sovereign GitLab ERP system deployment
How we set up a private GitLab instance to guarantee our business continuity
Business continuity
management is
one of our lines of business. How do we practice it ourselves? The
brain
of the company is an enterprise resource planning (ERP)
system that is in charge of all business related activity, such as the
publication of this blog post.
Based on prior experience we chose GitLab. Not the service gitlab.com but an installation of the open source software in our private network.
To ensure data security, we run our own LDAP service and private key
infrastructure (PKI). The containers of all services are built
based on GitLab pipelines and schedules in closed gitlab-runner
containers running in an isolated environment. Situational awareness
is provided by a centralised logging service. This is also how the
GitLab service that we offer to our customers works.
To us, GitLab is not only a handy web interface to the git version
control system. Business planning, management, standardised operating
procedures, customer projects and the source code of the entire
infrastructure are managed in the system. It allows us to plan, review
and test all changes before their deployment in production. The system
of continuous integration and continuous delivery (CI/CD) keeps among
others our VPN gateway and name servers up
to date, no matter if we desire to change something or some security
updates or certificates need to be managed or renewed.
For improved security, the validity period of our private certificates is limited to a few days. Should a certificate be compromised during its validity period, we can promptly cancel it via our CRL and OCSP based revocation service.
The validity period of our public certificates is determined by their issuer, Let’s Encrypt. Also the validity periods of public certificates will be gradually shortened from months to weeks. A short validity period will basically require automated renewal.
If we were the registrar of the domain example.com and our customer
wanted to set up a service https://download.example.com, we would
update our name servers and issue a public or private certificate for
the name, by making the changes in GitLab.
Containerised services
To ensure data security and manageability we have composed the services into containers that communicate with each others via carefully defined TCP/IP ports. The containers make it easy to keep the services up to date and to duplicate or move them between servers. GitLab pipelines and schedules control how, when and where the containers will be updated, tested and deployed in production.
Our service containers are based on a trusted up-to-date base image, often alpine:latest, which we adjust to the task at hand based on rules that are maintained under version control.
The basic idea is that GitLab will regularly rebuild the entire contents of each container as directed by the version control system and run functional tests on them before deploying them in a production environment. For instance, should our web server be subjected to a successful attack, the attacker might temporarily modify the contents. An automated update (or a manually applied one) would rectify the situation, possibly applying an update that would close the security hole that enabled the attack.
The core of our GitLab service consists of containers built on GitLab Docker images and a Postgres image. Supplementary services include an LDAP service and an email service.
Cleaning up containers
A downside of our method of work involving frequently rebuilt service containers is the huge storage consumption. Even though the containers can normally be rebuilt fairly quickly, we back up our container repository. This allows us to restore various services even when a degraded Internet connection would make some operating system distributions unavailable.
To save space, we have developed a staggered clean up service that retains all containers that are currently needed in production, as well as containers that are under active development. The removal must not simply depend on the last modification date, because the renewal schedule varies between services. Furthermore, the contents could remain unchanged for a long time, in case neither the base image nor our adjustments have changed since the previous update.
Recovery objectives and level of risk
There are two aspects to recovering a service from a disruption. The recovery point objective (RPO) defines the maximum acceptable interval during which data is lost. The recovery time objective (RTO) is the duration from the disruption to restoring the service.
Almost all our company data except for email folders and securely stored private keys is stored in the GitLab service. Thus, it is paramount to ensure its availability.
A failure of our ERP system does not directly impact the production services, which are deployed in completely different environments in accordance with our layered data security architecture. Only if a production service happened to be disrupted at the same time with the ERP system, could it take somewhat longer to restore the observable service.
Our recovery time objective is rather lax, perhaps a few working days. That would be the worst-case time needed for setting up the ERP system in a new piece of hardware at a different location, for example in the event of a fire.
Our recovery point objective is the last full backup, which currently
is one week. We could lose a few working days of data. The most
important of them could be restored manually based on emails sent by
the system as well as on git repositories located on personal
devices.
The recovery time and the level of risk of any system can be drastically reduced by implementing a high availability service. In case of a disruption, we could quickly fail over to a hot standby service. However, it is difficult to implement and test real-time replication between geographically distributed systems.
Backed up snapshots for disaster recovery
For maximum operational reliability, we avoid incremental backups. Our full backups comprise a logical backup of the database contents. Even though the persistent state of GitLab is divided into loosely coupled components, it is possible to create a consistent snapshot without shutting down the system.
Restoring a logical snapshot of a database can be more time-intensive than restoring a physical backup, which would include readily built index data structures. With the logical form, the data volume is smaller and a possible internal corruption of the database will be healed when restoring the backup. The logical form is also more reliable in the event of system upgrades that might change the physical format.
It is good to start any system update right after a backup has been made. In that way, should any problems arise, the system can be rolled back to the old data content and software version. An update could involve some storage format changes that will prevent a return to an earlier version. Particularly large updates should be tested carefully in a separate staging environment before changing the production environment.
We limit the space consumed by backups. Before initiating the backup procedure, we clean up the containers, which will shortly block the execution of GitLab pipelines. Furthermore, we employ the Grandfather-Father-Son backup rotation scheme: A few latest backed up snapshots are always retained. Of backups older than a couple of months, we retain one per month for up to a couple of years. Of older backups, we retain one per year.
Our justification for this pruning of old backups is that it is a rare
event to delete data, usually mandated by data retention policy.
Between backup snapshots, records are typically only being added. If
any content that is under version control is modified or deleted, the
new contents will be added in its own changeset and the entire change
history will remain available in the git repository. Only the
removal of a repository would be final—and extremely rare.
Testing recovery
An attempt to restore a backed up snapshot must be made regularly, before a disaster occurs. Otherwise, we could find out the hard way that the latest usable backup is very old.
Because our data is encrypted not only during transit but also at rest, restoring a backup involves encryption key management. The keys are of course kept separate from encrypted snapshots, because storing them together would make the encryption useless.
We maintain records of regular dry runs of disaster recovery from a remote backup.
Conclusion
A company that operates in a regulated business field must guarantee its data security, data privacy and business continuity. The service level objective will be reached via an ERP system that supports standardised operating procedures and automated processes. Nowadays, it is possible to achieve all this in a vendor independent manner with the help of open source software.