This commit is contained in:
2026-09-17 12:44:32 +08:00
parent a004d6110b
commit 4728a37c78
7 changed files with 1131 additions and 0 deletions

View File

@@ -0,0 +1,181 @@
## Security
### Access Control
**AWS Account Access Review**
现在来说一下 SOC2 audit 在我们实际的工作当中是怎么样来执行的。先开始第一个:security 里
面的 access control。讲的例子主要是我们要怎么样来管理 AWS account 的 access control。因为我作为整个 cloud service delivery 的 manager,所以我管理了差不多有 23 个 AWS account。我是这些 AWS account 的 account owner。所以我会每个季度进行一次 quarterly review 来检查这些 account 里面访问的用户的一些权限,包括检查一些已经离开公司的人员是否还有访问权限。先我们在 AWS account 里面事先会设计不同的 IAM 角色,然后给这些角色赋予不同程度的权限。比如 operation team 的 member 的权限比较大,因为他要在 AWS 上面去操作很多资源来进行一些升级,包括一些网络的修改等等。
其次我们还有一些 DevOps 团队。这些团队其实有些是 Network Team,有一些是 CCOE Team,他负责的是一些操作系统的 AMI的更新。所以他们的权限相对来说比 operations 要稍微小一点,主要集中在他们所对应的相关领域里面。我们还设计了一些 read-only 的权限,这个主要是给到一些 architect。他们可能要在我们的环境里面去抓取一些相关信息,但是他们不允许在我们的 AWS account 里面做一些修改,所以我们会设定一些 read-only 的政策。当然,这个里面权限设定会比较复杂,还有其他各种各样的权限都是根据不同的角色来设定的。你的操作方法呢是:我们每三个月会从整个 AWS account 的 primary owner 那边去拿到一个 account 的 access list。因为我们的 AWS account 访问是和 company 的 single sign-on 绑定的,所以我们是通过 email authentication 来登录 AWS account。所以这些信息会在更高级别的这个 IT 团队里面有,访问具体某个 AWS account 的 user 的一个 list。我拿到了过去3个月的访问记录以后,我会逐个地对这些人员进行分类:
- operation 团队的人
- devops 的
- 其他的 architect group 的
分好类以后我会检查他们所对应的权限。如果出现一些人员我并不认识或者是我并不组织的,我会发给相应的 devops 的 manager 以及 architect 的 group 的 manager,请他们帮忙来进行确认。 包括我自己的 operation 团队,我也会检查所有人员是否是当前在职人员。如果有一些人员已经离职了,但是他的访问权限还出现在这个 access list 里面,那我们就必须采取行动立即停止这些访问权限。 这个动作称之为“calibrate access write”。
当我完成了每三个月一次的 access review 以后我会做一个 summary report,然后把这个 report 发给我的 director、manager,甚至高级别的 high-level manager,来进行 sign-off,来确认我们完成了这样的一个动作。
在这个过程当中所有最初的原始人员的名单、我进行 calibrate 的 report,以及我去跟其他的 manager 进行确认的沟通邮件,都会被记录下来作为 evidence,来应对后续的 SOC2 audit。
**Restrict to access product source code repository**
还有一个项目也是会定期来做检查的。这个的目的也是从安全的角度考虑。在我们现在的这个组织里面我们是不允许 operation 的管理员有权限去访问整个产品的 source code repository 的。
operation 它应该不应该去放到 product 的 source code 不能做任何的 code commit 来修改产品里面的任何一些逻辑,包括一些安全方面的东西。它是不能涉及到这个 source code 的。所以在这个地方我们也有额外的 check:我们会检查我们的 GitLab 的权限,确保我们的 operation engineer 没有任何权限去访问我们的 source code 的这些 repository。
地方呢,我们会找到相应的 product team 的 owner、manager,作为这些 code repository 的 owner。我们会要求他能够提供这么一个名单,然后我们会来检查我们的 operating ID 是否在这个名单里面。正常的应该不在那个名单里面。
如果我们发现有些工程师是有权限可以访问的,我们会做一些 cataloger 去把这些权限可能拿掉。同样的整个过程当中所有的记录,原始的记录包括 cataloger 的记录,以及后续的一些 summary 的 report,都会保存下来,也会给high level manager sign off。
**Risk Assessment**
**Vulnerlubility Management**
这里介绍的是一个关于 vulnerability 的管理。我的介绍的 case 是我们的商业应用在云上的环境下面,我们定期会有 security 的 scan。其中就有 Qualys 的 scan 和 Prisma Defender 的一些 scan。
Qualys 的 scan 主要是针对我们 Cloud Application Runtime 环境上面的操作系统,比如 Linux,包括会有哪些 risk 和一些漏洞,这些都会被 Qualys 进行扫描出来。我们是这样来进行管理的。
作为整个 account 的 owner,我会定期收到系统发出的一些 policy report。这些 report 的内容会非常大,涉及整个 OS 里面的很多方方面面的一些 vulnerability。我们会对这些问题进行 filtering。首先我们会根据里面一些问题的 severity 来进行 filtering,来分析哪些是 OS 级别的。我们会结合我们另外一个 branch,那是我们整个 cloud central excellence team 提供的持续的 OS 级别的 Linux AMI hardening。
最早我们一开始使用的是 AWS 原生的 Linux hardening。后期因为各种客户的需要,包括我们有一些 specific 的 security requirement,我们就开始 adopt CCOE 提供的标准 Linux AMI。它发布的周期差不多是每个月发布一个新的版本,包含最新的 patch。我们会在收到他们测试过的 AMI 之后来进行 project plan。 除了通过AMI的升级能够修复一些OS级别的问题之外,我们还会定义我们还会 filtering 一些是不是通过OS升级来发现的那些问题。
针对这些问题我们就需要去做一些额外的动作。有可能我们会要去通知 product team 升级某些 library 的版本。如果一些旧的版本包含了一些 vulnerability 的问题的话,我们需要去联系研发团队、开发团队去更新这样的 library。
最终我们会定义一些 plan:哪些问题我们会在下一个版本进行修复。当我们升级完 AMI,包括开发团队也提供了相应的一些产品补丁之后,我还会针对这个最新的 policy 来和之前的 review 进行一次比较。看看哪些问题可以标注为已经修复了,哪些问题其实还是没有修复呢。如果有些问题并没有完全修复,我们会 highlight 出来,然后再进一步制定一些新的计划。计划还是根据整个扫描出来的问题的 severity 来进行下一步的新的计划。
比如之前Qualy扫描出来的有些问题是由于 Kubernetes 的版本太低造成的。我们就需要计划在下一个版本中对 AWS 上面的 EKS 的版本进行升级来解决这个问题。
像类似的这种问题我们都会进行检查,并把所有的发现、filtering、review,包括后续 fix、testing、sign off,这些所有的内容整个过程全部都记录下来,以便以后进行后续的 software audit。
## Availbility
- **Backups**
针对我们 SaaS application 的 backup,我们基本上是 90% 依赖于 AWS 原生的一些 cloud-native backup feature。
比如说我们的 application 首先是基于 AWS EKS 整个进行 Kubernetes 容器化部署的。所以在备份过程当中整个 cloud application 的备份,我们并不是要备份所有这些庞大的 images,因为这些 images 通过 R&D 团队每个版本发布到 GitHub 上面去的。
唯一我们需要备份的其实是 Kubernetes 的一些 configuration files。这样的话我们可以通过 configuration files(那些 yaml 文件)能够快速地把整个 web application 整个容器化,包括它的 pod 分布和 work node 的分布。我们可以很快地按照这个 configuration 还原出来。
所以我们只要备份一些 configuration files 就可以了。数据库RDS和数据存储EFS这方面我们是依赖 AWS backup。 这个 AWS backup,我们是会制定一系列的 backup plan,包括了整个备份的频率。我们每 6 个小时会对 RDS/EFS 的整个数据库和数据存储进行一次备份。这个备份不仅仅是备份到当前的 region;同时我们会通过一些额外的脚本来实现把 RDS 能够备份到一个remote region:
AWS Oregon -> AWS North Viginia
AWS Frankfult -> AWS Ireland
然后我们整个数据保留 7 天。这个 backup 是根据我们对整个 disaster recover 的一个 commitment 来做到的。RPO/RTO 我们在 SaaS 的一个 service description 里面是提到的,我们会把数据的影响控制在大概 6 个小时之内。
- **Processing Capacity**
- **Replication**
- **数据复制**:在多个系统/地域间同步数据副本,确保即使主系统故障,数据仍可被访问
- **实时或准实时的备份机制**,避免单点故障导致数据丢失
- **恢复时间目标(RTO)的支撑**:通过即时可用的副本,快速切换到备用系统
我们每 6 个小时会对 RDS 的整个数据库进行一次备份。这个备份不仅仅是备份到当前的 region;同时我们会通过一些额外的脚本来实现把 RDS 能够备份到一个remote region:
AWS Oregon -> AWS North Viginia
AWS Frankfult -> AWS Ireland
到的这个方案基本上还是属于一个 cold backup、远备份的方案,因为是考虑到一些成本的原因。其实在这个基础上我们还有一套更快速能够缩短整个 RTO 时间的预案,我们称之为一个热备份方案。
那个方案基本上是会把 RDS 和 EFS 这些 snapshot 在另一个 Viginia 的 AWS 环境里面定时地直接恢复到这个数据库里面,而不是以 snapshot 的形式存在。这样做虽然我们并不是说真正去启一套 runtime 的 instance,但是它可以大大缩短我们在另外一个 region 恢复整个 cloud 的状态所需的时间。这样的话这个成本会相对来说有所提高。
在这个方面我们会根据客户一些不同的要求来进行取舍。
- **Multi-location Strategies**
- **Business Continuity Planning and Testing**
- **Disaster Recovery Planning and Testing**
我们一年会进行两次DR testing,我们称之为 disaster recovery integrity testing。测试一次是轻量级的,不影响生产环境的,只是验证数据完整性的测试。那另一次就相对来说更复杂一点,它是一个完全DR的测试。
是针对我们的 backup 的数据。这里面包含了 device 的 configuration file 和我们的 RDS 数据的备份。根据这些内容在我们一个 backup 环境里面恢复备份数据。这个 backup 环境跟生产环境在同一个 region 但是不影响正常环境。
我们去利用备份数据恢复一套账,看整个恢复需要多长时间,但是这个恢复不切换生产环境。
第二次测试我们会 test 这个 remote 的 backup。因为我刚才提到了,在备份过程中我们会把数据备份到另外一个 remote region。我们利用这些 remote region 的 backup data 去 restore 整个 instance,然后这个是要配合在生产环境。
我们基本上是要在周末做停机切换:会把生产环境停掉,然后拿最新的 backup 的 delta 来恢复 remote region 的一个 instance。恢复好了以后我们会把整个生产环境的流量切换到新的 remote region 的环境下面,再恢复整个数据的流量,让客户能够使用。基本上是在一个星期之后,我们把数据反向恢复到我们原来的旧生产环境当中,再做一次服务压力测试,给它切换回去。
这个对我们的要求会比较高。
抵押的标准要求是按照我们在 SAS service description 里面给用户的承诺:RPO 和 RTO。 Recovery product objective 是 6 小时,recovery time objective 是 24 小时。这个策略也是根据这个标准来执行的。
## Confidentialty
- **Confidential Information Classification**
- **Confidential Information in Non-Production Environments**
针对这一块的 software audit,我们主要是用来证明我们在测试环境方面有专门的 staging,用于测试 application 的 upgrade,包括一些 patch、hotfix 的 deployment,以及我们在调整 cloud infrastructure 的结构上面的一些部署。包括一些自动化的测试都是在 staging 环境上面进行测试的。
这个的目的就是要测试我们在 staging 的环境上面没有用到大数据上面的 custom、confidential 的一些数据。这个主要的证明方法其实就是我们提供相应的 staging form 上面一些 tenant 的名称,包括做一些实际的演示,来证明我们仅仅只用到了一些测试的数据,并没有用到真实的客户的数据;也绝对不会包含 PII 个人信息的一些内容。
- **Data Deletion and Removal Practices**
Customer Exit Process:
## Introduction
When a SaaS customer decides to leave, it's crucial to handle the transition smoothly and professionally to ensure a positive experience, which can impact future business opportunities and the company’s reputation. This document describes the main processes and actions regarding customer exits.
## Service Description about Service Decomission
Service Decommissioning
Upon expiration or termination of the SaaS Order Term, Micro Focus may disable all Customer access to
SaaS, and Customer shall promptly return to Micro Focus (or at Micro Focus’s request destroy) any
Micro Focus materials.
Micro Focus will make available to Customer any SaaS Data in Micro Focus’ possession in the format
generally provided by Micro Focus. The target timeframe is set forth below in Termination Data
Retrieval Period SLO. After such time, Micro Focus shall have no obligation to maintain or provide any
such data, which will be deleted in the ordinary course.
### **Communication and Coordination**
- **Notify Relevant Teams**: Inform all relevant internal teams (support, billing, account management, cloud service etc.) about the customer's decision.
- **Designate a Point of Contact**: Assign a single point of contact to manage the transition and ensure all queries are addressed promptly. Usually it's the CSM.
### **Data Management**
- **Data Backup and Export**: Ensure the customer can export their data easily. Provide assistance if necessary.
- **Data Deletion**: Plan for secure deletion of the customer’s data from your servers after a certain period, in compliance with data protection regulations and your data retention policy.
- **Data Access Period**: Provide a clear timeline for how long their data will remain accessible after service termination.
### **Security and Compliance**
- **Revoke Access**: Ensure all user accounts associated with the customer are disabled and access to the system is revoked.
- **Compliance Check**: Ensure that the termination process complies with all relevant legal and regulatory requirements, such as GDPR or CCPA.
## Detailed Steps for customer exit
### Customer to submit service request to trigger customer exit project
The customer needs to submit a service request in PCS to start the customer exit process. All related communication will be still handled in PCS until all the tasks are done and close the user account in PCS.
- In the request, the customer needs to clarify the following specific needs:
Whether they wish to export existing ESM/SMAX transaction data?
What's the expected date customer want all tenant data to be emptied out completely?
What's the exact date Opentext to commit all relevant date (including backup data) will be cleaned out completely?
What's the exact data to close PCS support channel?
### Assist with data export
- What’s the suggestion to customer to export data?
- SMAX
- SMAX Offer customer to use OData export to export data
- Cloud Ops team can help to use existing OOTB OData export script to export SMAX transaction data per tenant
- CMS/HCMX/OO
- Not support by now
- PCS data
- No Support by now
### Plan data deletion
- Notification to customer to notify when we will terminate the tenant and delete all data
- Cloud Ops will handle such notification from PCS.
- Scope of data deletion
- Tenant data/user data/ account data (in BO)
- Inactive PCS entitlement 
- Data retention- farm level data retention is only 7 days. After 7 days customer data will permanently removed from Cloud environment
## Processor Integrity
[[Major Incident Management Process]]
[[Cloud Change Management Process]]
## Privacy
**Data Controller**
在这个方面我们实际的操作过程当中是的确有这样的一个需求的。我们管理的所有环境当中有两个比较特殊的环境:
- 在美国的一个 instance,因为它是要符合 FedRAMP 整个标准规范的。因此他对所有能够接触到数据的工程师有这样的要求:只有在美国当地的美国公民才能去 touch 这个环境里面的所有数据。因此我们在做数据控制方面会特意规避这样的要求。我们只允许美国的 operation 工程师能够操作、访问以及维护这一套环境上面所有的数据。包括在其他国家的中国、印度以及在欧洲的工程师都没有权限去 touch 这个数据。
- 在欧洲的环境,它是要符合欧盟的一些规范,包括一些 GDPR 的要求。同样地,只允许欧洲的工程师访问;其他 region 的工程师都不能访问那几个特定的环节。
[[GDPR]]
[[FedRAMP Basics Understanding Federal Cloud Security Standards]]